nafnaf.ai

ONE API. EVERY GPU.

THE NETWORK

One market. Every useful GPU.

Distributed NVIDIA and Apple silicon, verified and exposed through one control plane—without capacity tickets or cloud contracts.

DRAG THE PLANET TO EXPLORE

MANAGED INFERENCE

Requests in. Tokens out.

Use an OpenAI-compatible endpoint. We handle placement, health, failover, and price-aware routing behind the API.

No machines to manage

JOBS

Spot for flexibility. Protected for continuity.

Run checkpointable work on Spot capacity, or choose Protected execution when a marketplace interruption is not acceptable.

One workload spec, two execution contracts

HARDWARE

CUDA speed. Unified-memory scale.

NVIDIA ranges from RTX 3090 to B300. Apple starts at M3 Ultra, with native MLX and Metal runtimes for large-memory models.

24 GB to 512 GB accelerator memory

FOR PROVIDERS

Your hardware. Your terms.

Set availability and price floors. Serve inference, Spot Jobs, or Protected Jobs—and earn when your machines do useful work.

Three ways to run

PRODUCT

01

MANAGED

Inference

Send OpenAI-compatible requests. nafnaf routes each one to healthy capacity and streams the response back.

02

INTERRUPTIBLE

Spot Jobs

Run checkpointable batch workloads on the lowest-priced available hardware, with retries when capacity moves.

03

PROTECTED

Protected Jobs

Keep a GPU allocated for work that cannot be marketplace-preempted, with clear runtime and budget controls.

HARDWARE

Two architectures. One control plane.

Workloads declare what they need. nafnaf selects compatible capacity without pretending every accelerator is interchangeable.

NVIDIA · CUDA

RTX 3090 to B300

  • 24 GB or more VRAM
  • Managed inference and containerized jobs
  • Consumer cards through datacenter clusters

APPLE · METAL

M3 Ultra and newer

  • 96–512 GB unified memory
  • Native MLX, Metal, and approved runtimes
  • Large-model inference and memory-heavy jobs

How the market works

PROTOCOL

01

Verify

Providers connect eligible hardware. We benchmark performance, memory, thermals, storage, and network quality.

02

Route

Builders choose a workload contract. The control plane matches architecture, region, price, and reliability.

03

Settle

Usage is metered consistently across providers. Builders pay for execution; providers earn for useful compute.

FOR BUILDERS

Run work, not infrastructure

  • One API for inference and scheduled jobs
  • Spot or Protected execution per workload
  • Explicit hardware, region, and budget constraints
REQUEST COMPUTE →

FOR PROVIDERS

Put idle compute to work

  • Connect NVIDIA GPUs or an eligible Mac Studio
  • Control availability, workload modes, and price floors
  • Earn across inference, Spot, and Protected execution
LIST YOUR HARDWARE →

EARLY ACCESS

Be first on the grid.

Join the founding cohort of builders and compute providers. We’ll send your invite as soon as access opens.

No spam. Just launch updates and your access invite.