olmo-eval: An evaluation workbench for the model development loop
Imported from official source
olmo-eval is an open evaluation workbench that helps model developers add, run, and analyze benchmarks across changing LLM checkpoints, extending OLMES from final-score reproducibility into the day-to-day model development loop.
This version
- Version
- 1 of 1
- Recorded
- September 20, 2026 19:52
- Change
- Initial
- Content hash
37101c797f7dc490e27912488825959f- All versions
- Revision history