Scalable Question-Centric Text-to-Image Evaluation: Reliable Ranking, Fine-Grained Diagnosis, and Cost-Aware Routing
What changed: The proposed framework reports finer-grained ranking and diagnosis than aggregate scores, measures joint success across requirements, and supports cost-aware, training-free routing with 21.3% lower GPU-seconds per metric point than the cited ERNIE estimate.