LFM2.5-Encoders for Fast Long-Context Inference on CPU

Huggingface··Submitted by Mads Kristian Nylund
Open Source AIAI ToolsAI Infrastructure

LFM2.5-Encoders are optimized for efficient CPU inference, supporting up to 8,192 tokens with slow latency growth, and outperform ModernBERT-base in long-context tasks. They are built from the LFM decoder backbone, featuring bidirectional attention and non-causal convolutions, and are available on Hugging Face with demos for zero-shot prompt routing, policy linting, and spell checking. The models offer cost-effective and fast inference for tasks like intent routing and PII detection.

Read Article

More from Huggingface

Related Articles