Gliff will give you access to all your remote Omarchy boxes from a single app with tabs per server that streams over SSH using GPU acceleration. Incredibly useful! Will be part of RS 4.5 🤘 github.com/omacom/gliff
Tengo dudas, va muchísimo más rápido por DDR5 que por NVMe, aunque sea Gen5 en RAID 0. La latencia de la RAM está en nanosegundos; la del SSD, en microsegundos, y eso se nota en tokens/s.
If you’re someone who has 1 or more expensive GPUs and you want to run bigger models without dropping 5 digits to get more cards
1. Buy M2 > PCIe gen 5 raid0 adapter
2. Buy 4x 1TB Samsung gen 5 NVMe
You can get 60 GB/s on these fast enough to move data in and out of GPUs
Local IA is lit ¡Gran trabajo, Palmer! Los arreglos en el serve para que el agentic no se caiga (sobre todo tool calls y long-horizon) se notan claro en los números, y lograr eso sin tocar los weights es muy sólido. ¿Cuánto sacaría aproximadamente en tokens/s en una RTX 3060 de
Bonsai 2 27b PTQ1_0 from @PrismML is now smarter than ever!
12gb GPU, 100tok/s, and agentic workflow failures fixed!
After a week of experiments, I managed to find the two places that agentic work broke down in the model weights. Instead of adjusting the weights directly, I