Skip to main content
Back to Blog
Liquid AI's LFM2.5-VL-3B: A Game-Changer for On-Device AI Vision and Tool Integration
news

Liquid AI's LFM2.5-VL-3B: A Game-Changer for On-Device AI Vision and Tool Integration

Liquid AI introduces a powerful 3B vision-language model that runs locally on devices, revolutionizing screen understanding and tool calling capabilities.

3 min read

Liquid AI Releases LFM2.5-VL-3B: Compact Power Meets On-Device Intelligence

Liquid AI has just unveiled LFM2.5-VL-3B, a 3.1-billion parameter vision-language model designed specifically for on-device deployment. This release marks a significant milestone in making advanced AI capabilities accessible without relying on cloud servers or constant internet connectivity. With the model fitting into approximately 3 GB of memory, it's poised to transform how users interact with visual content and execute complex tasks locally.

What Makes This Model Stand Out?

The LFM2.5-VL-3B introduces three core capabilities that set it apart in the crowded vision-language model landscape:

  • Screen Understanding: The model achieves an impressive 80.7 average score on ScreenSpot-v2, demonstrating exceptional ability to read and interpret user interface elements. This is crucial for automation tasks, accessibility features, and intelligent assistants that need to understand what's displayed on your device.
  • Object Grounding: RefCOCO performance jumped dramatically from 57.1 to 87.9, meaning the model can now accurately identify and locate objects within images with remarkable precision. This improvement is substantial and opens doors for more reliable visual search and object-focused applications.
  • Function Calling: New to Liquid AI's vision-language line, function calling enables the model to determine which tools or actions to execute based on visual input. ToolSandbox performance surged from 26.4 to 59.5, more than doubling its previous capability.

Why On-Device Matters for Users

The shift toward on-device AI processing represents a fundamental change in how users will experience artificial intelligence. Rather than sending sensitive screenshots, documents, or personal images to cloud servers, the LFM2.5-VL-3B processes everything locally. This approach delivers several tangible benefits:

Privacy and Security: Your visual data never leaves your device, addressing growing concerns about data privacy and corporate surveillance. This is especially critical for users working with confidential documents, medical records, or sensitive information.

Speed and Reliability: With decoding speeds of 228 tokens per second on an Apple M5 Max, responses are nearly instantaneous. There's no waiting for cloud API calls or dealing with network latency issues. This responsiveness makes the model practical for real-time applications.

Cost Efficiency: Eliminating cloud processing reduces ongoing operational costs. Users and developers won't need expensive API subscriptions for tasks that can run on their own hardware.

Broader Implications for the AI Landscape

This release signals an important industry trend: AI capabilities are becoming increasingly sophisticated while simultaneously becoming more compact and accessible. The LFM2.5-VL-3B demonstrates that you don't need enormous models with billions of parameters to achieve strong performance on practical tasks.

For developers and enterprises, this means building smarter applications without heavy infrastructure requirements. For consumers, it means AI assistants that respect privacy while delivering intelligent features. For accessibility advocates, it enables local processing of visual information, making assistive technologies more responsive and reliable.

The addition of function calling to the vision-language model family is particularly noteworthy. It bridges the gap between understanding what's on screen and taking meaningful action—whether that's launching applications, filling forms, or executing workflows.

The Bottom Line

Liquid AI's LFM2.5-VL-3B represents a meaningful advancement in democratizing advanced AI capabilities. By combining powerful vision-language understanding with practical on-device deployment, the model addresses real user needs around privacy, speed, and functionality. As the AI industry continues evolving, expect more models optimized for local processing—making sophisticated artificial intelligence a feature of your device, not a service from the cloud.

Tags

vision-language modelson-device AILFM2.5-VL-3BLiquid AIscreen understanding
    Liquid AI's LFM2.5-VL-3B: A Game-Changer for… | aitoolfinder.ai