Show HN: Lumabri – Run Moe Models on a P2P Swarm with Colibri

Lumabri enables distributed inference of large mixture-of-experts (MoE) AI models across peer-to-peer networks, allowing organizations to run enterprise-scale models without centralized GPU infrastructure or expensive upfront data transfers. This P2P approach democratizes AI deployment by leveraging heterogeneous hardware (GPUs, CPUs, SSDs) across an organization, with intelligent caching that optimizes subsequent inference requests and maintains service continuity even if the primary node goes offline. For IT leaders, this represents a fundamental shift in AI infrastructure strategy—reducing capital expenditure on specialized hardware, improving fault tolerance, and enabling distributed inference at scale without traditional high-availability requirements.

Hacker News3 min read
Read full article
Show HN: Lumabri – Run Moe Models on a P2P Swarm with Colibri

Read the full story at Hacker News →