We're releasing APEX-Accounting, built with Ramp to test whether AI agents can complete real accounting work to professional standards. The benchmark includes 160 tasks across 10 simulated companies. Claude Fable 5 leads at 56.4%, but the headline result is not which model wins. It's how rarely any model succeeds consistently.


