Machine learning
AI, WITH THE RECEIPTS.
Know the score.
See the context.
Explore published model and agent results. Compare the test setup, the measured performance and the reported cost—then read the original source.
01 / THE BENCHMARK DESK
A score needs a setting.
Pick a test suite. Search the runs. Add up to three to your comparison.
- Filtered rank 1
Opus 5.5
Source run ↗max64.8%±3.11 pp
95% CI half-width - Filtered rank 2
Sonnet 5.5
Source run ↗max61.8%±2.94 pp
95% CI half-width - Filtered rank 3
GPT-6.1 Sol
Source run ↗max58.2%±3.11 pp
95% CI half-width - Filtered rank 4
Fable 5.1
Source run ↗max57.9%±3.76 pp
95% CI half-width - Filtered rank 5
Opus 5
Source run ↗max51.8%±3.39 pp
95% CI half-width - Filtered rank 6
GPT-6 Sol
Source run ↗max49.4%±3.23 pp
95% CI half-width - Filtered rank 7
Fable 5
Source run ↗max44.5%±3.85 pp
95% CI half-width - Filtered rank 8
GLM-5.3
Source run ↗max41.8%±3.23 pp
95% CI half-width - Filtered rank 9
GPT-5.6 Sol
Source run ↗max37.3%±3.78 pp
95% CI half-width - Filtered rank 10
GLM-5.3-Flash
Source run ↗none35.8%±3.54 pp
95% CI half-width
No runs match these filters.
Try another model name, task cohort or setup.
Ranks are within the filtered cohort. USD costs are publisher-reported evaluation totals, not API prices. A run’s agent, settings and source date matter. An unlisted model has no imported result, not a score of zero.
02 / SIDE BY SIDE
Your comparison desk.
Same suite, same task cohort. Different agent settings remain visible.
Start with a real question.
Which setup scored higher? What did the evaluation cost? Add two or three runs to see the details together.
03 / THE READING ROOM
Research, from the source.
Free, licensed RSS metadata from arXiv and Creative Commons. Original writing stays with its publisher.
A feed you can take with you.
New research announcements and open-source posts, with attribution.
Artificial intelligence
The Answer Is Not the Argument ↗
Machine learning
PEACE: Covariant learning of nonadiabatic manifolds with parity-resolved Hamiltonians ↗
Artificial intelligence / Language models / Machine learning
Training Parallel Speculative Draft Models by Directly Minimizing Expected Decoding Rounds ↗
Artificial intelligence
Justice After Identity: Large Language Models and the View from Everywhere ↗
Artificial intelligence
One-Slide Calibration of Pathology Foundation Models ↗
Artificial intelligence / Machine learning
Insights Generator: Systematic Corpus-Level Trace Diagnostics for LLM Agents ↗
Artificial intelligence
World Potential Model: Pretrained World Knowledge as Progress Potentials ↗
Machine learning
Derivative Gaussian Processes on a Two-Direction Budget ↗
No matching posts.
Try a broader title, author or topic.
arXiv items are research announcements and may be preprints, not peer-reviewed findings. We merge duplicate links across topics. arXiv dates are feed announcement dates; Creative Commons dates come from its Atom feed. Benchmarks refresh daily; feeds are checked hourly.
04 / READ THE FINE PRINT
Evidence before the headline.
Keep the tests separate.
Terminal tasks, code editing and refactoring measure different work. Even different versions of the same benchmark are separate comparisons.
Check what was tested.
The model, agent, reasoning effort, task count and evaluation date all affect the score. A stronger result on one suite is not a universal recommendation.
Follow the provenance.
Every run links to a revision of its publisher’s data. We show refresh status, licences and historical coverage, and retain earlier data if a refresh fails.