Accelerating vision-language models with LFM2.5-VL-DSpark

Imported from official source

Research

AI Classified by Officially

  • Faster inference: decode speedups up to 3.13x on device and 2.66x on an H100, with end-to-end gains up to 2.62x and 2.27x.
  • Small memory cost: the drafter adds 280M parameters, 8.9% on top of the 3B target
  • Day-one support: LFM-compatible DSpark integrations for llama.cpp, MLX-VLM, and SGLang
  • How does speculative decoding work for VLMs

    The vision drafter uses the same architecture as our text LFM2.5-DSpark drafters: it captures the target model's hidden states at a fixed set of tapped layers and conditions on them to draft a block of k candidate tokens. Image patches and text tokens are projected into a shared representation before those layers, so the drafter operates on hidden-state vectors of identical dimensionality regardless of input modality. The inference algorithm is therefore unchanged from the text models.

    We follow the DSpark recipe with a mixture of vision-language SFT data, weighted toward the workloads we expect the model to serve. Based on ablations across 3, 4, and 5 layers, the draft model is a simplified attention-only drafter with 4 layers and a block size of 9. We ran 10 epochs on the final mixture and measured acceptance after each, which improved with additional training tokens before reaching diminishing returns. At inference time, we recommend a block size of 8 or 9 depending on the hardware.

    The resulting drafter has approximately 280M parameters and increases the deployed model’s parameter count by just 8.9%.

    The DSpark draft model for LFM2.5-VL-3B ships with day-one support for llama.cpp, MLX-VLM, and SGLang.

    We measure both on-device inference and GPU inference. Both configurations use a DSpark block size of 8 and are evaluated on six diverse vision-based tasks, including general VQA, text VQA, image captioning, chart VQA, complex reasoning, and multi-turn conversation, following the MMSpec benchmark.

    This is an extract. The publication continues at the source.

    Read the original at the source: https://huggingface.co/blog/LiquidAI/lfm2-5-vl-dspark

    Officially imported this from Hugging Face’s own source and shows an extract. If you work there, claiming the profile and verifying the domain lets you choose to show the full text here.

    Provenance

    Organization
    Hugging Face — imported from official source
    Official source
    https://huggingface.co/blog/feed.xml RSS
    Imported
    September 24, 2026 15:00
    Versions
    1 recorded
    Identity
    https://huggingface.co/blog/LiquidAI/lfm2-5-vl-dspark

    Officially records where a publication came from, not whether it is true. Imported records are reproduced from an organization's own official source.