Chat with 100+ AI models privately on your device. Run Llama, Gemma, DeepSeek, Mistral offline — or connect to Claude, GPT-4, Gemini. Free and open source.
No reviews yetBe the first to leave a review for Fluent AI
Maker
📌
I started building this because I was frustrated: every AI app either locked great models behind a paywall, sent my conversations to the cloud, or only worked on one platform. I wanted something that actually respected my privacy and worked offline.
FluentAI now supports 3 local inference engines under the hood:
llama.cpp (GGUF) — the widest model compatibility, runs on every platform
LiteRT-LM — unlocks Android NPU acceleration on Snapdragon devices, giving you 2× faster inference than CPU
MLX — Apple Silicon native, so your M-series Mac or iPhone 16 Pro runs models at full speed
The result: you can run Gemma 4, Llama 3, or Qwen 3.5 completely offline, at full speed, with no API key, no subscription, and no data leaving your device.
For those who do want cloud AI, everything is there too — Claude 3.7, GPT-4o, Gemini 2.5, and 200+ models via OpenRouter, all with BYOK so you pay OpenAI/Anthropic directly (not us).
The biggest thing I've been working on recently: AI Agents — give FluentAI a goal and it'll plan and execute multi-step tasks using tools (web search, file system, calendar, MCP servers). It runs entirely on-device if you have a capable enough model loaded.
I'd love to hear what features matter most to you, and which models/platforms you're hoping to see supported. Happy to answer anything! 🙏