A no-code LLM playground to vibe-check and compare quality, performance, and cost at once across a wide selection of open-source and proprietary LLMs: Claude, Gemini, Mistral AI models, Open AI models, Llama 2, Phi-2, etc.
Hello Product Hunt community! 🚀
We're very proud to introduce the Airtrain.ai LLM Playground, a no-code tool to prompt many open-source and proprietary LLMs at once: Claude, Gemma, GPT-4, Llama 2, Gemini, Phi-2, Mistral models, and more. Compare quality, cost, and performance.
We built this playground to help AI enthusiasts and practitioners of all stripes easily “vibe check” popular LLMs.
Key features include:
📌 Prompt multiple models at once
📌 18 models supported (8 open-source, 10 proprietary)
📌 Inference metrics (i/o token counts, throughput, inference cost)
📌 Persisted sessions (review and resume previous chat sessions)
We'd love for you to try it out and share your feedback with us.
Feel free to ask any questions, and we'll be more than happy to answer them. Thanks so much for your support, and we hope you enjoy using the LLM Playground! ✨
Report
Emmanuel and team, congratulations on the launch! I have to say, it is so very cool to be able to query Claude 3, GPT-4, and Gemini all at once. What inspired you to build this?
@joy_larkin We heard from a lot of users that they wanted to try other models than the OpenAI ones but they did not know where to start.
It starts with evaluation, and evaluation starts with a vibe check. Running a handful of prompts through various models to build an intuition about their relative quality, cost, and performance, before moving to batch evaluation across an entire test dataset (which we also offer :)
That’s very interesting, I really like the concept of “vibe checking”. Is it possible to upload many samples I may have or run multiple times at once for a prompt? Since due to temperature responses vary right
@rchaves Yes, to do this, go back to the task menu (New task in the top bar) and select "Evaluate Models". You can upload a CSV or JSONL file of up to 10k examples and configure the models you want to test and the metrics you care about.
See docs here: https://docs.airtrain.ai/docs/ba...
@will_jacob Indeed Opus is quite expensive! Luckily Sonnet is still close to GPT-4 level but at a more reasonable price. It'll be interesting to have a chance to play with Haiku when Anthropic releases it, since that should be much more affordable.
Hi @emmanuel_turlay I love Airtrain.ai's comprehensive playground for LLM comparison!
It's a brilliant tool for developers and AI enthusiasts like me :)
Have you considered integrating a community-driven feature where users can share their findings or best practices with specific models? I think this would help with the platform's utility and have shared insights.
Airtrain.ai LLM Playground
Airtrain.ai LLM Playground
LangWatch
Airtrain.ai LLM Playground
Startup Digest
Airtrain.ai LLM Playground
Airtrain.ai LLM Playground
Airtrain.ai LLM Playground
Airtrain.ai LLM Playground
Airtrain.ai LLM Playground
Kniru: AI-Powered Finance