Liquid AI’s Fastest Vision Model Yet
Liquid AI brings vision AI to devices, Google enables sign-language translation, and Microsoft’s MAI-Image-2.6 reaches No. 2 in image generation.
This week in AI, the spotlight is on on-device vision AI, accessible human-computer interaction, and next-generation image generation. From Liquid AI bringing powerful vision-language capabilities to smartphones, to Google DeepMind enabling real-time sign-language translation, and Microsoft pushing its image model to No. 2 on the Arena leaderboard, the industry is moving toward AI that runs locally, understands people better, and delivers more capable creative tools.
Liquid AI is bringing fast vision AI to edge devices, with LFM2.5-VL-3B delivering strong screen, UI, grounding, and multi-image understanding while running in roughly 3 GB of memory.
Google DeepMind is making sign-language interaction more accessible, with SL2T translating more than 50 sign languages into streaming text and bringing sign-to-text capabilities directly to smartphones.
Microsoft is pushing image generation forward, with MAI-Image-2.6 reaching No. 2 on the Arena leaderboard and improving text rendering, photorealism, 3D imagery, and creative control.
Together, these developments highlight a broader shift in AI: powerful models are becoming smaller, more accessible, and more capable across vision, communication, and creative workflows.
Liquid AI Launches LFM2.5-VL-3B for Fast, On-Device Vision AI
Liquid AI has released LFM2.5-VL-3B, a 3.1-billion-parameter vision-language model designed to deliver strong visual understanding without the heavy compute requirements of larger models. The model improves significantly over its predecessor in screen and UI understanding, function calling, image grounding, and multi-image analysis, while remaining a non-reasoning model optimized for low latency. It scores 80.7 on ScreenSpot-v2, raises RefCOCO grounding precision from 57.1 to 87.9, and improves multi-image benchmarks such as BLINK from 50.2 to 61.5. Liquid AI also scaled its training to around 34 trillion tokens, expanded the tokenizer to 128K vocabulary, and increased vision pretraining data fourfold. What makes the release particularly notable is its edge performance: it can run in roughly 3 GB of memory, reaching 228 tokens/sec on an Apple M5 Max, 116 tokens/sec on AMD Ryzen AI Max+ 395, and even 20 tokens/sec on a Galaxy S26 Ultra. The model supports llama.cpp, MLX, vLLM, SGLang, and ONNX, making it suitable for everything from document and screen understanding to local AI assistants and visual agents.
Google DeepMind Brings Sign Language Translation Directly to Smartphone
Google DeepMind has introduced Sign-Language-to-Text (SL2T), a multilingual AI model designed to translate sign language into streaming text and bring sign-language interaction into everyday consumer devices. The technology now powers sign-to-text dictation in Gboard and Live Transcribe on Pixel 11, initially supporting American Sign Language (ASL) to English, with additional devices and languages planned. Instead of treating sign language as a simple sequence of hand gestures, SL2T understands the coordinated movements of the hands, arms, torso, head, and face, recognizing that sign languages have their own grammars and linguistic structures. The model was trained on more than 100,000 hours of data covering over 50 sign languages, with about a quarter of the data coming from ASL. For privacy, an on-device model converts the camera input into body-pose landmarks, meaning the original video can be discarded before the coordinates are sent for translation. SL2T reportedly achieves a 70 BLEURT zero-shot score on the FLEURS-ASL benchmark, while also being optimized for real-world challenges such as low latency, one-handed signing, left-handed users, and avoiding hallucinations when someone is not signing. Users can sign to search the web, draft messages or documents, respond during conversations, or even interact with Gemini without needing to type.
Microsoft’s MAI-Image-2.6 Climbs to No. 2 on the Arena Leaderboard
Microsoft has launched MAI-Image-2.6, its latest text-to-image model, which debuted at No. 2 on the Arena text-to-image leaderboard, ahead of leading models from Google, Meta, and xAI. The model gained 79 Elo points over MAI-Image-2.5, with improvements across every measured Arena category and a particularly strong 91-point increase in text rendering. Microsoft says MAI-Image-2.6 delivers better portraits and 3D imagery, stronger photorealism, and more polished commercial visuals for product, branding, and cinematic applications. Beyond image quality, the model introduces improvements in multi-reference handling, grounding, reasoning control, format selection, and resolution control, giving creators more control over the final output. The release continues Microsoft’s progression from MAI-Image-1 through MAI-Image-2 and 2.5, with the company positioning 2.6 as another step toward models that are more useful for professional creative workflows. MAI-Image-2.6 is currently available to try through Arena, with availability planned for Microsoft’s Playground, Foundry, and other products.
Hand Picked Video
In this video, we’ll look at Caveman, a Claude Code skill by JuliusBrussee designed to reduce token usage and make LLM outputs more concise and efficient.
Top AI Products from this week
Caveman - One command wraps Claude Code, Codex, Hermes, and more with a local proxy that compresses logs, tool output, and files before every provider call. In a pinned 54-run benchmark: 33.2% fewer input tokens with 18/18 correctness checks.
Mem Agent - An important deliverable, or a pizza place saved for that someday trip to Italy—Mem Agent keeps track of what you tell it, plus the todos living inside your notes and meetings, sharply following up so it actually happens.
FluidDocs CLI - FluidDocs CLI turns your prompts into interactive documents that answer questions and report back. You see who opened it, how far they got, and what they asked. Pitch decks, proposals, reports, board updates.
WebBrain - Your browser, your models, your data. WebBrain is a free, open-source AI browser agent for Chromium browsers and Firefox. Run it locally with llama.cpp and most queries cost nothing your data never leaves your device.
Execlave - Execlave is an AI Agent Governance and Enforcement platform (runtime AMP) that sits between autonomous agents and your real systems, enforcing policy before every action instead of after incidents.
Qencode MCP - Qencode lets AI assistants transcode, analyze, edit, optimize, and deliver video using natural language, powered by a cloud video processing platform.
This week in AI
Dots3 Raises the Bar - Dots3 is a powerful reasoning model built for complex tasks, delivering stronger problem-solving, coding, and agentic capabilities with improved efficiency.
Sarvam Just Made AI Voice Agents Available to Everyone - Sarvam has opened its Voice Agents platform to everyone, enabling natural, context-aware AI voice conversations at scale for businesses and developers.
Seedream 5.0 Pro Can Finally Understand Design - ByteDance’s new AI image model goes beyond generation with precise editing, complex infographic creation, realistic visuals, layer separation, multi-image fusion, and multilingual support.
Qwen Just Dropped a 2.4 Trillion Parameter Beast - Qwen3.8 brings 2.4T total parameters, 95B activated, stronger coding and agentic abilities, and up to 1M-token context—now available as an open model.
MAGI-2 Just Raised the Bar for AI Video - Sand.ai’s MAGI-2 Preview is a 114B-parameter audio-video model that activates only 6B parameters per token, generating 10-second 1080p videos with synchronized sound from text or images.
Paper Of the day
MARC is an open-source multi-agent framework designed to make clinical AI reasoning more structured, transparent, and reliable. Instead of relying on one LLM to handle everything, MARC assigns specialized agents to extract information, reason through it, generate answers, and evaluate results. Its Decomposer can also turn a plain-language task into a complete agent pipeline without manual prompt engineering. The framework supports configurable workflows, model swapping, RAG, API-based deployment, and local CPU inference, making it adaptable for biomedical Q&A, radiology reporting, and other clinical tasks.
Read this whole paper 👉 here




