Your Agent Aced the Task. Will It Do It Again?

Imported from official source

Research

AI Classified by Officially

That is embarrassing onstage. In production, it is a reliability problem: a workflow that succeeded once may fail the next time a user makes the same request. For mission-critical work, such as reconciling a financial transaction or checking a contract for an obligation, that can be a showstopper.

Most benchmarks hide this variability behind an average. On AppWorld, a ReAct agent using GPT-4.1 succeeded on 77.4% of runs across five repetitions. But it succeeded in all five runs for only 53.0% of tasks — a 24.4-point consistency gap.

Most benchmarks report the first number. We built a way to measure the second — and improve it.

In an earlier post, we introduced ALTK-Evolve — a system that turns an agent's own past trajectories into reusable guidelines, distilled automatically and injected back at inference time. It measurably improves task success, but those results only asked the average-case question too. This post introduces consistency guidelines, a new guideline type in altk-evolve built on top of a diagnostic tool we call the Consistency Analyzer, that targets this gap directly.

  • Accuracy hides an unreliability problem. A ReAct agent (GPT-4.1 on AppWorld test_normal) that succeeds 77.4% of the time on average succeeds on all 5 repeated runs for only 53.0% of tasks — a 24.4-point consistency gap. On hard tasks it reaches 30 points.
  • We built a diagnostic for exactly this. The Consistency Analyzer resamples an agent's own recorded trajectory to find flip-prone decision points — steps where the model was one token-sample away from doing something different. It needs one trace and no ground truth — it resamples each decision point in that trace with a single call requesting k completions (k=5 by default), rather than re-running the task end-to-end.
  • Turning that diagnosis into guidelines halves the gap — from 24.4pp to 12.0pp (same-task Pass⁵ +16.0pp, similar-task +13.0pp), without costing anything in average accuracy.
  • This is an extract. The publication continues at the source.

    Read the original at the source: https://huggingface.co/blog/ibm-research/altk-evolve-consistency

    Officially imported this from Hugging Face’s own source and shows an extract. If you work there, claiming the profile and verifying the domain lets you choose to show the full text here.

    Provenance

    Organization
    Hugging Face — imported from official source
    Official source
    https://huggingface.co/blog/feed.xml RSS
    Imported
    September 15, 2026 19:08
    Versions
    1 recorded
    Identity
    https://huggingface.co/blog/ibm-research/altk-evolve-consistency

    Officially records where a publication came from, not whether it is true. Imported records are reproduced from an organization's own official source.