01
MANAGED
Inference
Send OpenAI-compatible requests. nafnaf routes each one to healthy capacity and streams the response back.
ONE API. EVERY GPU.
THE NETWORK
Distributed NVIDIA and Apple silicon, verified and exposed through one control plane—without capacity tickets or cloud contracts.
DRAG THE PLANET TO EXPLORE
MANAGED INFERENCE
Use an OpenAI-compatible endpoint. We handle placement, health, failover, and price-aware routing behind the API.
No machines to manage
JOBS
Run checkpointable work on Spot capacity, or choose Protected execution when a marketplace interruption is not acceptable.
One workload spec, two execution contracts
HARDWARE
NVIDIA ranges from RTX 3090 to B300. Apple starts at M3 Ultra, with native MLX and Metal runtimes for large-memory models.
24 GB to 512 GB accelerator memory
FOR PROVIDERS
Set availability and price floors. Serve inference, Spot Jobs, or Protected Jobs—and earn when your machines do useful work.
PRODUCT
01
MANAGED
Send OpenAI-compatible requests. nafnaf routes each one to healthy capacity and streams the response back.
02
INTERRUPTIBLE
Run checkpointable batch workloads on the lowest-priced available hardware, with retries when capacity moves.
03
PROTECTED
Keep a GPU allocated for work that cannot be marketplace-preempted, with clear runtime and budget controls.
HARDWARE
Workloads declare what they need. nafnaf selects compatible capacity without pretending every accelerator is interchangeable.
NVIDIA · CUDA
APPLE · METAL
PROTOCOL
01
Providers connect eligible hardware. We benchmark performance, memory, thermals, storage, and network quality.
02
Builders choose a workload contract. The control plane matches architecture, region, price, and reliability.
03
Usage is metered consistently across providers. Builders pay for execution; providers earn for useful compute.
FOR BUILDERS
FOR PROVIDERS
EARLY ACCESS
Join the founding cohort of builders and compute providers. We’ll send your invite as soon as access opens.