Technology
The performance layer behind private AI inference.
MoonMath.ai builds the most efficient, privacy-centric AI inference endpoint.
We are a small team of mathematicians and engineers building fast, private AI inference through low-level algorithms, systems engineering, and hardware-aware optimization.
PostsAll posts
SGLangAug 5, 2026
Zro ♥ SGLang
Read post →ResearchJun 23, 2026
HyperQuant: A Rate-Distortion-Optimal Quantization Pipeline
Read post →AMDJun 17, 2026
A Fast Attention Kernel for MI300X, Written in HIP, Not Assembly
Read post →AMDAug 5, 2026
A faster MLA decode kernel for Kimi-K2.7-Code on MI300X
Read post →Stack
Compression, kernels, and hardware sit under the API.
Agent layer
Claude Code, Codex, Cursor, Cline, Opencode, Hermes
API layer
OpenAI-compatible requests plus Anthropic-compatible Messages
Privacy layer
Multi-region inference with zero request retention by default
Compression layer
HyperQuant-style compression for weights, KV cache, and long-context efficiency
Kernel layer
Custom attention and serving kernels for model-specific throughput
Hardware layer
AMD, NVIDIA, and Google TPU deployment paths
