Chart understanding and visual reasoning.
Evaluates multimodal models on realistic chart understanding using 2,323 hand-curated arXiv charts, with descriptive questions (basic elements) and reasoning questions (synthesis across complex elements).
Model
Score
Muse Spark 1.1xHigh
86.4%
Qwen 3.8-MaxxHigh
Score on the original open-source benchmark: 94.1% / Score on this Mercor-extended benchmark: 82.9%
GPT-5.6 SolMax • Pro
Score on the original open-source benchmark: 92.1% / Score on this Mercor-extended benchmark: 82.9%
GPT-5.4xHigh
Score on the original open-source benchmark: 93.4% / Score on this Mercor-extended benchmark: 82.7%
GPT-5.5xHigh
Score on the original open-source benchmark: 94.6% / Score on this Mercor-extended benchmark: 82.7%
APEX NEWSLETTER
New benchmarks, leaderboard shifts, and research from the APEX team.
By subscribing you agree to receive updates from Mercor.Unsubscribe anytime.