Allen Institute for AI

allenai.org

Imported from official source

Allen Institute for AI — publications from its own official source.

Type
Company
Scope
US · national
Website
allenai.org
Feed
Atom

Publications 25

  1. Open-sourcing AstaBrief, the fast report-generation model in Asta

    We’re releasing AstaBrief, an 8B open-weights model for generating cited scientific reports, available in Asta’s Fast mode or to download and run on your own infrastructure.

    AI
  2. Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs

    Olmo-core 3 introduces a redesigned, fully open training stack for efficiently scaling mixture-of-experts models into the trillion-parameter range.

    AI
  3. What a crowdsourced game revealed about steering Olmo 3

    A crowdsourced game built on Olmo 3 showed how people can exploit unexpected model behaviors to stress-test prosocial AI evaluations—and how open access to a model’s internals can help researchers ...

    AI
  4. Teaching future scientists to interrogate AI tools for scientific discovery

    University of Washington students put Ai2’s AutoDiscovery to the test, showing how AI can surface promising scientific leads while making human judgment, domain expertise, and rigorous validation m...

    AI
  5. How Goodfire used Ai2’s open post-training stack to trace unwanted model behavior

    Goodfire used Ai2’s fully open post-training stack to predict LLM behavioral changes, trace unwanted model behavior back to individual training examples, and test targeted fixes without sacrificing...

    AI
  6. The hard parts of AI-assisted science

    At an Ai2 event marking our expanded collaboration with Providence Swedish, researchers explored the hardest problems in AI-assisted science: keeping systems steerable, grounded in human judgment a...

    AI
  7. BenchMIRT: What are LLM benchmarks actually measuring?

    BenchMIRT is a new method for auditing LLM benchmarks question by question, revealing which capabilities they actually measure and helping researchers build smaller, more focused, and easier-to-int...

    AI
  8. Ai2 and Providence Swedish Cancer Institute partner to advance AI-assisted scientific discovery

    Ai2 and Providence Swedish Cancer Institute are expanding their collaboration after AutoDiscovery helped researchers uncover and validate a promising new immune signal in invasive lobular breast ca...

    AI
  9. How researchers adapted Dolma for better Thai language models

    Thai researchers adapted Ai2’s open Dolma toolkit to build Mangosteen, a 47-billion-token Thai corpus that filters low-quality web data while maintaining or improving model performance and strength...

    AI
  10. How a Georgia Tech team used the open Olmo stack to trace social reasoning

    A Georgia Tech team used Ai2’s fully open Olmo stack to trace social reasoning back to the training data that shaped it, finding that dialogue-rich, interpersonal writing had an outsized influence ...

    AI
  11. When a model reads a drug's class from its name—not its knowledge

    Researchers used Olmo 3 and its open training data to show that models can infer a drug’s class from its name instead of knowing the specific medication, and traced that shortcut to how often drugs...

    AI
  12. TutorMoments: Do AI tutors know when to help and when to hold back?

    TutorMoments is an open, replay-based evaluation framework that tests whether AI tutors can recognize when to support a student and when to hold back and encourage deeper reasoning.

    AI
  13. Ai2 expands collaboration with Hugging Face to accelerate open science

    Ai2 is expanding its partnership with Hugging Face to give its growing portfolio of fully open models, datasets, benchmarks, and applications the storage, bandwidth, and integrations needed to reac...

    AI
  14. Tracing distinctive language in AI-written text

    Stony Brook researchers used our infini-gram engine to trace distinctive phrases in AI-generated writing back to existing sources, finding that top-selling self-published books on Amazon with subst...

    AI
  15. The OlmoEarth Platform: Geospatial inference at planetary scale

    How we built the OlmoEarth Platform to fine-tune geospatial models and run continent-scale satellite inference while managing massive data pipelines, distributed compute, and automatically recoveri...

  16. Who gets to understand AI?

    Why fully open models and research artifacts are essential to independent scrutiny, broader participation, and continued U.S. scientific leadership in AI.

    AI
  17. What building Shippy taught us about building agents

    Building Shippy taught us that reliable agents depend less on the model itself than on deterministic tools, explicit guardrails, isolated infrastructure, and evaluations grounded in real-world work...

    AI
  18. Modular LLMs at scale: how FlexOlmo is helping to pool national expertise without pooling sensitive data

    Danish Foundation Models is using FlexOlmo as the basis for FlexMoRE, a more efficient modular LLM architecture that lets institutions contribute specialized experts trained on sensitive or proprie...

    AI
  19. DiScoFormer: One transformer for density and score, across distributions

    DiScoFormer is a transformer-based density and score estimator that can infer both quantities from a finite sample in one forward pass, generalizing classical KDE while staying accurate in high-dim...

    AI
  20. Which tokens does a hybrid model predict better?

    New token-level analyses of Olmo 3 and Olmo Hybrid show that hybrid models predict meaning-bearing, context-dependent tokens better than transformers, while transformers retain an edge on verbatim ...

    AI
  21. MolmoMotion: Language-guided 3D motion forecasting

    MolmoMotion is an open, language-guided 3D motion forecasting model that predicts how object points will move in the future, enabling stronger motion prediction for robotics, video generation, and ...

    AI
  22. olmo-eval: An evaluation workbench for the model development loop

    olmo-eval is an open evaluation workbench that helps model developers add, run, and analyze benchmarks across changing LLM checkpoints, extending OLMES from final-score reproducibility into the day...

    AI
  23. Building accessibility tools on a truly open foundation

    PointCheck, an independent project, uses Molmo, MolmoWeb, and Olmo 3 to test web accessibility the way a keyboard user would—by navigating real pages and inspecting what's actually on screen.

  24. OlmoEarth v1.1: A more efficient family of models

    OlmoEarth v1.1 is a more efficient family of remote-sensing models that cuts compute costs by up to 3x while maintaining similar performance, making large-scale satellite mapping faster and cheaper...

    AI
  25. Introducing AIMIP: The AI weather and climate model intercomparison project

    AIMIP is a new open benchmark and dataset for evaluating AI climate models, showing they can match or beat conventional models on some historical climate metrics while still struggling to generaliz...

    AI

Officially records where a publication came from, not whether it is true. Imported records are reproduced from an organization's own official source.