Up to 3.2x Faster Inference with LFM2.5-DSpark

Imported from official source

Research

AI Classified by Officially

  • Faster inference: up to 3.18 throughput improvement on a GPU and up to 2.87x on-device.
  • Toward on-device agentic inference: cuts function-calling latency by 57% on average for LFM2.5-2.6B
  • Day-one support for llama.cpp and SGLang: LFM-compatible DSpark integration is open-sourced upstream
  • The decode phase in LLM inference is traditionally memory-bound. Most latency comes from streaming weights from DRAM into SRAM, not from intense computation. Speculative decoding addresses this by using a lightweight draft model to produce candidate tokens, then having the target model verify them all in a single forward pass, sharing the cost of loading the weights across all tokens we verify.

    Over the years, multiple approaches of speculation have been proposed, with the most prominent being EAGLE-3, DFlash, and, most recently, DSpark, which combines three components:

  • DFlash-style parallel backbone conditioned on the target model’s context features, producing hidden states for all draft tokens in a single forward pass.
  • A lightweight sequential head, modeled as a Markov chain between neighboring tokens, that adds inter-token dependency, raising the acceptance rate at later positions.
  • A confidence-scheduled verifier that predicts each token’s survival probability and prunes low-confidence suffixes when verification would cost more than it saves.
  • We follow the DSpark recipe with a larger and more diverse data mix covering SFT, chat, code, and function-calling data. Based on our ablations, the first versions of the draft models are simplified attention-only draft models, with 5 layers and a block of 9. For each draft model, we ran 15 epochs on the entire dataset and selected the epoch with the highest acceptance rate rather than the lowest loss.

    The resulting draft models are relatively small, with each around ~300M parameters.

    This is an extract. The publication continues at the source.

    Read the original at the source: https://huggingface.co/blog/LiquidAI/lfm25-dspark

    Officially imported this from Hugging Face’s own source and shows an extract. If you work there, claiming the profile and verifying the domain lets you choose to show the full text here.

    Provenance

    Organization
    Hugging Face — imported from official source
    Official source
    https://huggingface.co/blog/feed.xml RSS
    Imported
    September 15, 2026 19:08
    Versions
    1 recorded
    Identity
    https://huggingface.co/blog/LiquidAI/lfm25-dspark

    Officially records where a publication came from, not whether it is true. Imported records are reproduced from an organization's own official source.