MIT PhD Li Muyang Launches Nunchux AI to Accelerate Multimodal Inference
Nunchux AI has officially launched as a new player in the generative AI infrastructure space, founded by MIT doctoral graduate Li Muyang, Carnegie Mellon University associate professor Zhu Junyan, and researchers Yujun Lin and Zhekai Zhang. The company targets the growing industry demand for efficient, low-cost, and reliable multimodal inference as deployment scales rapidly. Backed by early-stage investments from Emergence Capital and E14 Fund, MIT Media Lab, Nunchux AI builds upon a decade of research into neural network compression and system optimization. Li’s academic work, particularly the SVDQuant framework for 4-bit weight and activation quantization, forms the technical foundation of the venture. This research materialized in Nunchaku, an open-source inference engine that drastically reduces memory footprint and latency, enabling high-resolution image generation on consumer hardware. The engine has secured significant developer adoption, recording nearly 4,000 GitHub stars and native compatibility with major open-source models and development frameworks. Commercially, the startup will debut Modelverse, a unified API platform integrating over thirty image, video, and virtual avatar models. The service offers two operational tiers: Radical Speed for ultra-low latency and Radical Value for cost optimization. Unlike existing inference aggregators that primarily manage GPU orchestration, Nunchux AI pursues vertical integration by combining custom quantization algorithms with optimized inference kernels. This approach delivers superior performance and lower overhead on identical hardware. The Nunchaku core will remain open-source, while commercial revenue will be driven by premium API access, enterprise-grade customization, and dedicated developer support. The venture extends the entrepreneurial legacy of MIT professor Han Song’s HAN Lab, which previously spawned companies such as DeePhi Tech, OmniML, Eigen AI, and Inco AI. Academic production continues in parallel, with recent publications including the FourTune paper explicitly listing the company as an affiliation. In a crowded inference market featuring competitors like fal.ai and Replicate, Nunchux AI differentiates itself through algorithmic efficiency rather than mere aggregation. The founding team maintains that reducing latency and deployment costs is critical for preserving creative velocity and democratizing access to advanced generative models, while strictly preserving output fidelity.
