Deploy local agents everywhere with LFM2.5-2.6B

LFM2.5-2.6B is built to power capable agents entirely on-device. It supports tool calling and multi-step workflows while staying small and fast enough for everyday hardware, from laptops to phones. This enables developers to deploy agents everywhere, keep data private on the device, and scale usage without a cloud inference bill.

  • Best-in-class agent: Competitive with models 4x larger on tool use, instruction following, and multi-step agentic tasks.
  • Agentic reinforcement learning: Trained inside the most popular agentic harnesses to improve compatibility.
  • Efficient inference: 220 tok/s on an Apple M5 Max and 113 tok/s on an AMD Ryzen CPU, in under 2.5 GB of memory.

lfm2_5_2_6b_evaluations



How we built a reliable agentic model for edge devices

LFM2.5-2.6B is pre-trained on ~34T tokens, with a mid-training phase that extends the context window to 128K. Post-training then turns the base model into an agent in four stages:

  1. Supervised fine-tuning (SFT): two rounds of SFT, weighted heavily toward agentic data like tool use, web search, and harness trajectories.
  2. Teacher specialization: train one specialist teacher per domain (math,

     

     

     

    To finish reading, please visit source site