← All Projects
MLXexo ringVRAMSovereign AIInferenceCost Optimization
Sovereign AI Inference Cluster
48GB Unified-Memory Sharded Ring — $26,400/mo Cloud Cost → $0 API (self-hosted)
$0
Cloud API Cost (self-hosted, from $26,400/mo)
Industry
AI Infrastructure / Enterprise ML
Duration
12 months
Roles
8 qualified
APKHVESOLDIAHCADMIDSEECN
The Problem
Enterprise inference at scale costs $26,400/month on cloud (OpenAI GPT-4 equivalent). Data sovereignty requirements (HIPAA, ITAR, EU GDPR) prohibit cloud inference for regulated use cases.
Approach
Designed and built sovereign inference cluster: M1 Max (32GB unified) + M4 Mini (16GB) = 48GB unified pool, sharded ring on a public fork of exo (upstream lineage). OpenAI-compatible API endpoint — drop-in replacement. Specialist lanes (code · reasoning · vision) served as 30B-class Q4. 55 commits ahead of upstream on the fork — clone and compare.
exoMLXApple Silicon M1/M4OpenAI-compat APIQdrant
Proof Metrics
$0Monthly API cost, self-hosted (from $26,400/mo)Cloud-API invoice vs local — internal billing, reproduce-on-request
389ms P50End-to-end inference latencyPrometheus-logged, self-hosted cluster — internal, reproduce-on-request
99.97%Uptime (12 months)Monitoring log — internal
55Commits ahead of upstream exogithub.com/mpuodziukas-labs/exo — clone & verify
Want this on your team?
This work maps to 8 qualifying roles. Remote, Q3 2026.