Product

A mill for inference silicon.

Dielet is an inference foundry. Performance engineers are agents that rewrite kernels. You buy the endpoint that those agents verified on your mix — serverless Turbo or a dedicated die.

Profile

Agents sit on your GPUs and flag the true bottleneck: kernel, batch, decode, or HBM.

Search

Candidates include fused kernels, FP8 passes, and new shard maps. Each is measured.

Verify

A candidate ships only if outputs match and the clock is lower on your mix.

Serve

OpenAI-compatible endpoints. Dedicated capacity in under a day.

Search space

What the mill actually tries.

Nothing ships because it looks faster on a blog. A candidate has to beat your recorded mix and keep the tokens.

Kernels and decode

Fused attention, KV layouts, and decode paths measured against the sequence lengths you actually serve.

Batching and shards

Batch shapes, FP8 passes, and shard maps. The mill keeps the map that won on your GPUs, not a generic recipe.

Weights

Open models on the rack.

Bring a public checkpoint or a LoRA. Named families we mill first:

  • Llama 3.1 70B
  • Llama 3.1 8B
  • DeepSeek V3
  • DeepSeek R1
  • Mistral Large

Serving

Two ways to take the die.

Hosted Turbo

Shared, optimized endpoints. Playground tokens on Spark; Forge when you want SSO and a 99.9% target.

Dedicated Die

Reserved H100 or H200, private VPC, 24-hour cutover. Fab if the mill has to run on-prem.

Teams

Who it is for

Builders

Keep your chat client. Point it at a Dielet base URL after the mill prints a verified die.

Operators

GPU-seconds, versions, and spend sit next to the same objects your engineers debug.

Security

SSO, audit logs, and regional storage. Ujjwal Kumar Singh signs the DPA.