Timbal helps teams turn AI prototypes into production systems. Build agents and workflows, connect them to your data, design interfaces, deploy, monitor, evaluate, and govern everything from one platform. Instead of assembling separate tools for retrieval, orchestration, UI, observability, and evals, Timbal gives you one core for shipping reliable AI applications.









Free Options
Launch Team / Built With

World model AI : Sharm SEKAISimulate your idea on a living world.
Promoted


Hey Product Hunt 👋🏻
I'm Martí, co-founder and CEO of Timbal AI.
In today's AI world, going from 0 to 1 it's easy fast and cheap, going from 1 to 100 is not.
Complex data, legacy software, messy folders and heavy compliance and cybersecurity requirements, that's the enterprise reality.
You can prototype an agent in an afternoon, but then the real work starts. You have to wire together a vector database, an orchestration framework, a UI tool, and an observability layer. Before you know it, you're maintaining a fragmented stack of vendors, and your app still breaks in ways you can't easily trace.
Timbal is the unified stack that closes that gap. Everything you need to go from idea to production lives in one place:
A Database Native to AI: Run vector, keyword, and relational searches in a single query, ensuring your agents are always grounded in your actual data.
Deterministic Workflows: A reliable runtime for agents with built-in observability, traceability, and evals. You will always know exactly why your AI made a specific decision.
Omnichannel Visual Builder: Turn your logic into real apps instantly. Ship directly to the web, WhatsApp, email, or voice.
Our core is open source (Python framework, NPM packages, and TypeScript SDK on GitHub). You can stay in the code, build visually, or seamlessly mix both.
Putting this in front of the PH community is the milestone we’ve been waiting for. We want your brutal feedback—tell us what you'd build, what's missing, and where you think we're wrong.
Let's chat in the comments!
Martí & the Timbal team
I like that you’re combining orchestration, deployment, observability, and evaluation instead of expecting teams to stitch together several different tools. I’m curious how opinionated the evaluation layer is—can teams bring their own eval datasets and metrics, or does Timbal encourage a particular workflow?
@amjad_shaik Good question Amjad! The short answer is that we give you the structure, you decide everything inside it.
Evals in Timbal run on a YAML-based test format, you define the runnable, the inputs, and what "correct" means for that specific case: output validators (type, pattern, semantic match), execution sequence validators (did it call the right tools in the right order), and timing thresholds. None of that is prescribed content, it's your dataset, your params, your definition of a passing result.
Where we're opinionated is the format and the mechanics: a consistent way to write, discover, and run evals via CLI, so they plug into CI/CD the same way across every project instead of everyone inventing their own testing setup. That consistency is what lets Composer also generate a starting set for you if you don't want to write from scratch, but it's optional scaffolding, not a constraint.
@inescastillo That clarifies it well. I especially like that the format is opinionated but the evaluation logic isn’t. I can see how that would make it easier to adopt across multiple projects without forcing everyone into the same testing criteria.
@pablo_veciana For Knowledge Bases specifically, ingestion goes through Timbal's own pipeline today, you add files via direct upload or a remote URL fetch, or upload structured/tabular data directly, and then you can query it (including raw SQL) once it's in. That's the native path, one hybrid store for retrieval.
But KBs aren't the only way to bring in your data. Agents can call external tools directly, including ones exposed via MCP, so if you've got a live data source you don't want to migrate, you connect it as a tool the agent calls at runtime instead of ingesting it into a KB. Timbal itself is also MCP-connectable the other direction, we support connecting to Claude, Cursor, VS Code, and other MCP clients directly.
In conclusion, native KB ingestion for data you want unified and queryable inside Timbal's retrieval layer, MCP/external tools for data you want to keep live and separate.
@pablo_veciana great question Pablo! Also to build on Inés' reply, once something's ingested, you're not stuck with however it auto-chunked. You can go in and edit individual chunks directly (what gets embedded vs. what gets shown), so that in the case that retrieval quality is off for one document you fix that document instead of having to re-tune the whole pipeline.
@ridhwikvinod On your specific pain point, rollback. Every workflow is a real git repo, commits, diffs, branches. A bad prompt change is just a commit you revert, same as reverting any other code change, no separate rollback UI to figure out under pressure.
On testing before production:,you define evals per step, so you're not just checking "did the final output get worse," you can pinpoint which exact node degraded. Composer can also scaffold a starting eval set for you if writing them from scratch isn't where you want to spend time.
For a 40-person team running this internally without a big platform team behind it, that combination (git-native rollback + step-level evals) is probably what saves you the most time day to day, not because either piece is fancy, but because you already know how to use them.
how do you manage version control across workflows. a simple visual history could help teams track every important change.
@darly_selby Great question! Every workflow in Timbal is a git repo under the hood full commits, diffs, branches, the works. You can clone it, review it on GitHub, and roll back exactly like any other codebase. Environments map to branches too, so promoting dev → prod is a merge, not a copy-paste.
Love the visual history idea though!
Congrats on the launch!!
That's a really cool project, I tried to create a simple agent and it really satisfied my expectations.
But what about token usage? isn't it expensive?
@yernururu Thanks, glad the agent held up! Fair question, and actually it tends to go the other way. ACE stops agents from wandering into extra tool calls, so token usage often drops instead of climbing, and overall ends up more economical 🙂
The "one stack" pitch is appealing — the amount of time spent stitching together separate tools for
retrieval, orchestration and observability adds up quickly. How does it handle model switching mid-workflow?
Curious whether you can swap between providers without rebuilding the whole pipeline.
@charles_mondal_phd Great question! And yes, that's genuinely just swapping a string, models are referenced as "provider/model" throughout, so switching Anthropic to OpenAI to Gemini mid-pipeline doesn't touch the rest of the logic. You can also chain providers as an automatic fallback (if one fails or times out, it tries the next) instead of hardcoding just one. No rebuild, no separate integration per provider.