Needle 2
A 14MB agentic LLM designed for tool calling and structured extraction on low-power devices like phones, wearables, and robots.
Category: AI Agents
Overview
Needle 2 is an open 45M-parameter model optimized for extreme efficiency. It runs as a compact 14MB binary using only 28MB of RAM, making it suitable for hardware with limited resources like microcontrollers and budget smartphones.
How to use it?
Developers can integrate the Needle engine into their applications to enable on-device tool calling and data extraction. It runs directly in environments like WebAssembly or on edge hardware to perform tasks like smart home control or data processing without needing a cloud connection.
Features
14MB binary size, 28MB RAM footprint, On-device tool calling, Structured data extraction, High-speed token decoding, Cross-platform edge compatibility


