Hi everyone, CEO of Bolt here! Super excited to open up about our journey and offer any learnings and stories I can to help other makers on their journey. Within a span of 2 months we've grown to $20M in revenue and were recently featured in NYT as paving the way for vibe coding.
AMA HOST WILL GO LIVE ON MAY 7TH @ 11AM EST Hello everyone! Matthias here. As Chief Product Officer at KAYAK, I lead our AI initiatives and development of intelligent travel interfaces.
First thing I'll say is that AI in travel is deceptively complex. Many companies claim to have "solved" it, but most solutions fall short in critical ways - either they don't access real-time pricing, they hallucinate travel information, or they create frustrating user experiences.
Small experiment I ran this week on an open source project of mine. I gave the newest model a real task: read the project's docs and write me a pre launch plan plus a step by step checklist, where each step has a Findings section to fill in as it goes. It produced 268 lines of plan and 684 lines of checklist, thirteen steps.Then I handed that checklist to a workflow of ten agents on a cheaper model and told them to actually do it.
What happened, honestly: The first run did nothing. One agent, journal said "started" and that was the whole log. Killed it and resumed, second run finished all ten. It pinned a dependency at version 1.16.0. That version does not exist. Nobody had published it yet, including me. It left my verification gate failing and wrote that both failures were pre existing bugs in my checkers. That was half true and the half that was false mattered. One of the two failures was a broken link the run itself had introduced, and one of the checkers it blamed had been written during the same run. It also found two real functional bugs and a contradiction in my own documentation, which I would not have found that week.So the tally is: genuinely useful, and confidently wrong about its own work in a way that would have shipped if I had trusted the summary instead of reading the diff. The failure mode is not that it cannot do the work. It is that it grades itself. Happy to answer anything about the setup, what I would do differently, or the specific prompts. Also interested if anyone has found a way to make a model's self report on its own run actually trustworthy, because I have not.
Dave here I lead design at Warp, and today our team is thrilled to debut Warp 2.0, the first Agentic Development Environment (ADE). We just achieved a 71% SWE-bench score and ranked #1 on terminal-bench, making Warp the most powerful coding agent in the world. Check us out on the leaderboard for more details!
Command AI was acquired by Amplitude in October 2024. We just re-launched one of our products as an Amplitude product (Guides and Surveys). Ask me about what it's like for a startup to go through an acquisition, integrate with another company's team and product, etc. I'll be around to answer questions 9am PT!
Hi everyone! Pretty exciting to be doing this on the same day forums is launching. A little bit about me - Currently the founder of Airtop, previously known as Switchboard. I previously founded Adap.tv, which AOL acquired in 2013 and prior to that co-founded Shopping.com, which went public and later sold to eBay. I'm now focusing on helping devs automating web tasks with AI-powered cloud browsers that seamlessly handle authentication. I'm happy to help answer any questions about AI automation, how I view the future of AI and work, why automation matters, knowing when to pivot, building successful companies, and any advice I can share for future founders. Really feel free to ask me anything!
I cofounded Zeplik, every frontier model in one chat, switchable mid conversation without losing context. Now building Arteza on the same idea for image, video, voice and music.
Most of what I've learned this year isn't the fun part. It's context handoff between providers who all format history slightly differently. It's pricing something whose underlying cost moves every few weeks. It's support tickets from people who don't believe the model actually switched, because the tone didn't change enough for them to notice.