Log inOpen app
Log inOpen app
kwindla
Daily
6,723 posts
kwindla profile banner
@kwindla

kwindla

Daily
@kwindla
Infrastructure and developer tools for real-time voice, video, and AI. @trydaily // ᓚᘏᗢ // @pipecat_ai
San Francisco, CA
machine-theory.com
Joined September 2008
3,948 Following
16.2K Followers
1 Subscription
RepliesRepliesRepostsRepostsMediaMediaArticlesArticles
Get the full app experience
Unlock more features and see what people are talking about right now.
Open X
  • Pinned
    @kwindla
    kwindla
    Daily
    @kwindla
    Aug 27
    Introducing PhoneLLM, an open model for voice agents. GPT 5.6 Terra performance on typical voice agent tasks at 1/3 the latency and 1/18 the cost. For voice agents, we need models that are both very low latency and very good at tool calling and instruction following. There's a
    00:00
    119
    250
    2.5K
    333K
  • @kwindla
    kwindla
    Daily
    @kwindla
    Oct 9
    More multi-model voice AI demos from Jon: this time a tiny bi-directional encoder model in the Pipecat pipeline that runs on CPU fast enough to be in the main voice loop. Fun fact: you can train a classifier model like this in less than 10 minutes on a modern GPU. (If you have
    @JonPTaylor
    Jon Taylor
    Daily
    @JonPTaylor
    Oct 9
    I’m obsessed with using tiny classifier models in Pipecat agents. This one, Audience, is a ~33M-param ONNX model built on Ettin that works out who each line in a scene is for. It uses names, descriptions, nearby items and where people are standing, and it can pick one character,
    00:00
    6
    8
    72
    3.8K
  • @kwindla
    kwindla
    Daily
    @kwindla
    Oct 8
    I really like the Bland Speech v3 voice model from Bland, and a lot of other people do too! It's currently the top realtime TTS model on the Design Arena Audio Realism ranking, which is a tournament-style benchmark in which actual people judge which model sounds more human. The
    @pipecat_ai
    Pipecat AI
    Daily
    @pipecat_ai
    Oct 8
    Sometimes, choosing a voice for your Pipecat agent is about more than just latency. It might be about personality, accent, or how convincingly a voice fits the situation. We explore voice realism with @usebland Speech v3, available in Pipecat as a maintained service. 📖
    00:00
    2
    3
    43
    3.4K
  • @kwindla
    kwindla
    Daily
    @kwindla
    Oct 8
    Add to the list of things I did not expect to be true when computers became super-intelligent: the single most important glue-layer tool (load-bearing, as the LLMs say) would be a terminal multiplexer written in 2007.
    9
    1
    13
    1.6K
  • @kwindla
    kwindla
    Daily
    @kwindla
    Oct 6
    This is awesome. Multi-character interactions in Unreal Engine, completely LLM-driven and unscripted. You can spin your own version of this up pretty easily. This is another Pipecat + Jev experiment from @JonPTaylor. Watch Jon walk around the game world, talk to the characters,
    00:00
    2
    4
    39
    2.3K
Edit with