LLMs and RAGs are powerful, but without a strong evaluation pipeline, you're flying blind. Here's how to systematically test and improve your GenAI Apps. 🧵 (1/6)
If models already stumble onto system exploits as a side effect of benchmark optimization, what happens when persistence is the explicit objective function?
What actually stops someone from using RL to reward root access, weight cloning, and anti-shutdown behavior directly?
Honestly pretty torn on open weight models after the OpenAI/Hugging Face situation.
On one hand, letting people fine tune away safety guardrails is a real risk. On the other, overly strict guardrails just tie the hands of security people when things go wrong.
How to keep AI spend flat while token usage grows exponentially: Not with friction and spend alerts. With better defaults, routing, and caching.
Better Defaults (not Usage Caps) – Engineers can choose any model they want, but defaults matter. We’re experimenting with defaulting
It’s pretty amazing to see how much the Claude Code/Codex default settings are designed to just burn tokens.
There are dozens of configuration changes you can make that will drop token consumption with zero impact on quality. Each one is about 1%-3% in cost savings, but together