Log inSign up
Log inSign up
Akshay πŸš€
21.4K posts
Akshay πŸš€ profile banner
@akshay_pachaar

Akshay πŸš€

@akshay_pachaar
Simplifying LLMs, AI Agents, RAG, and Machine Learning for you! β€’ Co-founder @dailydoseofds_β€’ BITS Pilani β€’ 3 Patents β€’ ex-AI Engineer @ LightningAI
Learn AI Engineering πŸ‘‰
join.dailydoseofds.com
Joined July 2012
502 Following
290.5K Followers
2 Subscriptions
RepliesRepliesRepostsRepostsMediaMediaArticlesArticles

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
TermsΒ·PrivacyΒ·CookiesΒ·AccessibilityΒ·Ads InfoΒ·Β© 2026 X Corp.
  • Pinned
    @akshay_pachaar
    Akshay πŸš€
    @akshay_pachaar
    Nov 14, 2023
    My lecture at MIT!✨ From Physics to Linear Algebra & Machine learning, I have learned a lot from MIT! Yesterday, I had the honour of delivering a guest lecture on The state of AI Engineering, exploring: - Prompt Engineering - Retrieval Augmented Generation. - Fine-Tuning
    134
  • @akshay_pachaar
    Akshay πŸš€
    @akshay_pachaar
    5h
    This inference engine runs LLMs 4x faster than llama.cpp: (Qwen3.5-9B at 92 tokens/s on 16GB Mac) I ran Qwen3.5 9B on a base M5 MacBook Pro using three local inference engines, one at a time. β†’ llama.cpp reached 22.0 tokens/s β†’ MLX reached 25.1 tokens/s β†’ Uzu reached 92.1
    00:00
    @akshay_pachaar
    Akshay πŸš€
    @akshay_pachaar
    21h
    Article cover image
    Article
    Karpathy’s Trick for Faster Local LLMs, finally has a proper solution
    In 2023, Karpathy explained why local LLMs are slow. When a model generates one stream of text on your own computer, the chip spends most of its time waiting for weights to arrive from memory while...
    36
  • @akshay_pachaar
    Akshay πŸš€
    @akshay_pachaar
    21h
    Article cover image
    Article
    Karpathy’s Trick for Faster Local LLMs, finally has a proper solution
    In 2023, Karpathy explained why local LLMs are slow. When a model generates one stream of text on your own computer, the chip spends most of its time waiting for weights to arrive from memory while...
    7
  • @akshay_pachaar
    Akshay πŸš€
    @akshay_pachaar
    Oct 7
    A faster, 100% open-source alternative to Jev. It’s called Laya. Both models solve the same problem. You give them unstructured state plus typed questions. Instead of generating a paragraph, they choose from answers you define and return probabilities your code can act on.
    @akshay_pachaar
    Akshay πŸš€
    @akshay_pachaar
    Sep 18
    Article cover image
    Article
    Jev Clearly Explained
    We have been using LLMs like a hammer for every AI problem, even simple decisions. Jev handles those decisions in milliseconds at a fraction of the cost. Let's understand how it works and where it...
    25
  • @akshay_pachaar
    Akshay πŸš€
    @akshay_pachaar
    Oct 7
    LLM Routing, clearly explained! (routing can cost more than not routing at all) The basic idea sounds simple: Send easy requests to cheaper models and reserve frontier models for difficult tasks. But production routing is rarely that straightforward. Every classification
    00:00
    @akshay_pachaar
    Akshay πŸš€
    @akshay_pachaar
    Sep 6
    Article cover image
    Article
    LLM Routing Can Cost More Than Not Routing
    A classifier reads the request, a cheap model handles the easy ones, and the frontier model handles the rest. That is the version of routing everyone knows, and it is also the version that breaks in...
    23
Edit with