Has anyone tried in organization to use self hosted llm models for agentic programming?
Im curious if it makes any sense. My organization spends fortune on tokens from US companies. I want to recommend something… I think that will be cheaper to use it on own machines instead…
mistral?
So self hosting is still not great.
The big problem is you can get large memory but slow prompt processing, which reduces your context window, or you can get semi-fast GPU with low memory, where you’re capped on models.
Sometimes I run pi agent in a container with Gemma 4 or Qwen 3.6, but even on strix halo after 60k tokens the quadratic slowdown is brutal.
We aren’t there yet for complex agaentic workflows locally, and it’s primarily a hardware issue.
Though innovations in performance are being shipped regularly, they’re incremental.
Is 128 GB of ram per unit enough for your organization’s use case? You could convince them to buy a Framework Desktop and then install an offline llm to it (ollama with Mistral, perhaps). Then you don’t have to rely on American companies or the environmental impact of data centers, and then after the startup cost, it’s free from then on.
Best of all, they can just be normal work computers when the bubble bursts.
I wish I could just say, “Convince your company not to use AI,” but I’m sure your higher-ups aren’t taking no for an answer.
Something like AIhorde as the foundation ?
self hosting an orphan crusher doesn’t sound like a meaningful improvement
Pi.dev with Qwen3.6 running on a modest 6GB GPU is actually working pretty well for me. For smallish well-scoped agentic code tasks.




