Open-model inference

Cut the die. Keep the tokens.

Point Dielet at an open checkpoint and the GPUs you already rent. Agents search kernels, batching, and decode layouts against recorded traffic, then serve the stack that measured faster — same silicon, identical outputs.

  • Same GPUs measured on your mix
  • Match first required before cutover
  • Open weights Llama, DeepSeek, Mistral
Llama 3.1 70B Llama 3.1 8B DeepSeek V3 DeepSeek R1 Mistral Large Open weights Same GPUs Llama 3.1 70B Llama 3.1 8B DeepSeek V3 DeepSeek R1 Mistral Large Open weights Same GPUs

Method

How Dielet ships

Point at a checkpoint. Let the mill search overnight. Swap the base URL when the die is cut.

01

Point at the checkpoint.

Bring an open weight or a LoRA. Name the GPU class you already rent. Dielet records a slice of your real traffic.

02

Let the mill run.

Agents search kernels, batching, and decode layouts against that mix. Every candidate is measured, not guessed.

03

Cut over.

Swap the base URL. Tokens keep their shape. Latency drops on the same silicon. Serverless or dedicated.

Stack

The mill, in four cuts

Performance engineers are agents. You buy the endpoint those agents verified on your mix.

Profile

Agents sit on your GPUs and flag the true bottleneck: kernel, batch, decode, or HBM.

Search

Candidates include fused kernels, FP8 passes, and new shard maps. Each is measured.

Verify

A candidate ships only if outputs match and the clock is lower on your mix.

Serve

OpenAI-compatible endpoints. Dedicated capacity in under a day.

Keep

What does not change.

Your GPUs.

Dielet measures on the class you already rent. No new silicon required for a first cutover.

Your tokens.

Cutover waits until outputs match. The mill is not allowed to ship a faster wrong answer.

Your client.

Keep the chat client. Point it at a Dielet base URL after the mill prints a verified die.

Start

A die on the rack this week.

Write ceo@dielet.website. You get a test token, a hosted Turbo endpoint, and a 30-minute walkthrough on the model you want to serve.

See pricing