Based AI Foundry

We build and serve open-weight models at scale.

Based is an AI foundry focused on quantization, fine-tuning, and inference infrastructure for frontier-scale systems. We make large models practical — reproducible, fast, and deployable on real hardware.

Compute online
118.8 TFLOPS
4× DGX Spark
200K ctx
TP4 · W4/W8

Full-stack model infrastructure

From raw checkpoints to production serving. We handle the hard parts of working with large open-weight models.

01

Quantization

W4/W8 and sub-4-bit (IQ1S) quantization of large models. Reproducible recipes that preserve quality while fitting on real hardware.

02

Fine-tuning

LoRA and overlay fine-tuning of frontier open-weight models — Kimi, GLM, and others — for specific domains and tasks.

03

Serving

Production inference on multi-node clusters with tensor parallelism (TP4), long context (200K), and low-latency decode. We publish the exact serving recipes.

04

Cluster ops

Operating DGX Spark and GB10-class hardware. Storage, orchestration, and reproducible builds across multi-node deployments.

05

Model releases

We publish our work openly — checkpoints, buckets, and recipes on Hugging Face, with full build scripts on GitHub.

06

Advisory

Hands-on guidance for teams adopting open-weight models — hardware selection, quantization strategy, and serving architecture.

Open, reproducible, public

Everything we build is published. Models and checkpoints on Hugging Face, build and serving recipes on GitHub.

Public work on GitHub

Reproducible builds and serving recipes, published openly. Everything is public and runnable.

View all on GitHub

Let's build something.

Tell us what you're working on. We'll get back to you within one business day.

Domainbasedllc.co