← arXiv

arXiv·3 min read

Statistical attribute alignment for black-box generative AI via output post-processing

It addresses how to enforce a desired attribute distribution when only black‑box queries to a generative model are available.

The authors consider a scenario where a user can repeatedly query a generative AI system without internal access, and must collect a batch of m outputs whose joint attribute distribution matches a prescribed target. This setting captures fairness constraints and representative synthetic data generation, where the attribute of interest could be gender, race, age, or any categorical variable.

To meet the alignment goal, they design exact and approximate post‑processing algorithms that select and possibly transform generated samples so that the expected number of queries needed is minimized. The methods are proven optimal in the limit as m grows large, and they operate purely on the observable outputs, leaving the underlying model untouched.

Experiments on a text‑to‑image generator and a geocoded persona synthesis task show that the post‑processing pipeline consistently brings the empirical attribute distribution closer to the target than baseline prompting techniques, confirming its practical benefit for statistical alignment.

TakeawayPost‑processing can efficiently enforce target attribute distributions in black‑box generative models, outperforming prompting alone.

Prodigy briefing — continue on the original for source material, discussion, and updates.

Read the paper ↗