Log inSign up
Log inSign up
Vinay Goel
Amplitude
343 posts
@vinaygo

Vinay Goel

Amplitude
@vinaygo
Staff AI Engineer. Leading Agent Analytics @Amplitude_HQ | Ex @includedhealth @internetarchive
San Francisco
Joined June 2009
480 Following
303 Followers
RepliesRepliesRepostsRepostsMediaMediaArticlesArticles

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • @vinaygo
    Vinay Goel
    Amplitude
    @vinaygo
    5h
    Thank you @digitalocean. Excited to be featured here and talk about what we’re building at @Amplitude_HQ
    @digitalocean
    DigitalOcean
    @digitalocean
    7h
    "Building an AI product goes beyond just selecting a model. It’s a systems problem.” @vinaygo of @Amplitude_HQ on how they built Wave, a proactive agent that finds what’s worth building and hands you ready-to-merge PRs. All the time. The data science agents behind it run on
    00:00
    2
  • @vinaygo
    Vinay Goel
    Amplitude
    @vinaygo
    8 Oct
    Your AI bill can drop 80%. Most teams get there by swapping in a cheaper model, and their users pay the difference. We measured it on our own agent: half the cost per user, 88% slower responses, 10% fewer messages. Here's how to cut the cost and keep the quality ↓
    00:00
    5
  • @vinaygo
    Vinay Goel
    Amplitude
    @vinaygo
    7 Oct
    Agent evals had two lanes: Code and LLM judges. Jev opened a third: decision models. OpenAI just entered it. OpenAI Decisions is now live alongside Jev in Agent Analytics! Dry-run it, then score production sessions at scale and chart the results alongside what users did next.
    @OpenAIDevs
    OpenAI Developers
    OpenAI
    @OpenAIDevs
    6 Oct
    01:11
    Let your app choose the right model, tool, or action in near real-time with Decisions API, now available to all developers in public beta. The Decisions API makes decisions up to 10x faster than GPT-6 Luna through the Responses API.
    8
  • @vinaygo
    Vinay Goel
    Amplitude
    @vinaygo
    7 Oct
    Step three: see if it worked for your users. An eval score isn't the outcome. 86% of CashBook's agent sessions passed the response-quality check. Looked fine. But users whose first answer went badly were retaining 15 points lower four weeks later, across all of CashBook, not
    @vinaygo
    Vinay Goel
    Amplitude
    @vinaygo
    6 Oct
    Step two: ship the fix. Topics show you where users are struggling. But before you fix anything, you need to know how big the problem is and what's causing it. A few HYBRD users said their AI coach sometimes needed a second nudge before it would run a tool. A few unlucky users,
    2
  • @vinaygo
    Vinay Goel
    Amplitude
    @vinaygo
    6 Oct
    Step two: ship the fix. Topics show you where users are struggling. But before you fix anything, you need to know how big the problem is and what's causing it. A few HYBRD users said their AI coach sometimes needed a second nudge before it would run a tool. A few unlucky users,
    @vinaygo
    Vinay Goel
    Amplitude
    @vinaygo
    5 Oct
    Step one: see the problem. That starts with a question most agent teams can't answer: what are users actually asking it? From the very first session, Agent Analytics groups conversations into topics and scores each one on task completion, friction, helpfulness, safety, and more.
    2
Edit with