Bonacci Foundry trains a domain model inside your perimeter and hands you weights you own outright. You do not need a corpus to start: a few hundred real examples is enough, and we build the rest. Start with a one week audit that will tell you whether you need a model at all.
One week corpus audit, free for now · model builds $15k to $120k · air-gap capable · weights delivered, not licensed
Usually that's about volume, not knowledge. We can solve for volume.
A few hundred real cases is enough to characterise a domain. Foundry expands that seed set into training volume synthetically, holding to its distribution so the generated data stays inside your world. It is then decontaminated against the eval suite, so the scores at the end still mean something.
We curate the corpus ourselves from public sources, licence filtered, and build against an evaluation suite we agree with you up front. This is how our own models were trained, including the finance model on Hugging Face.
Your documents, tickets, filings, logs, code. We measure duplication, PII and template collapse before a GPU hour is spent, and tell you in writing if what you have is too thin to justify the build.
# seed set you actually handed us seed examples 412 real deduplicated 397 quality filtered 361 # volume we generate from it synthetic expansion 361 → 84,000 grounded in seed distribution eval contamination clean # corpus we curate, when you have none public domain crawl 11.4B tokens licence filtered 8.9B total corpus 9.1B tokens # your 412 examples set the target. # they are not the training set.
A bad config burns the compute. We check everything that can't be undone before a GPU hour is spent.
One week, and we are not charging for it at the moment. We audit your data, measure tokenizer fit on your corpus, and train small models on your actual corpus to see where the curve goes. Most people estimate this from document counts. We measure it, and the report tells you what 9B buys you over a fine-tune in numbers from your own data.
Corpus checks measure duplication, PII and template collapse before compute is spent. Configuration checks block the specific mistakes that cost us those three models. Training then records a second curve alongside the loss, because this class of failure does not show up in the loss at all.
Before anything ships we open the weights up and look inside them: dead embedding rows, layer contribution, special-token health, then a smoke test on the serving stack you will actually run. You get weights you own outright, the eval harness, and the configs to reproduce the run.
# checked before a single GPU hour is billed P1 weight decay excludes norm gains BLOCK exclusion list is empty: every parameter decays, including the RMSNorm gains. invisible in loss. P2 checkpoint interval fits tenure BLOCK 10000 steps on spot capacity: every preemption restarts at zero P15 licence permits delivery terms pass P16 eval sets decontaminated pass P17 tokenizer fertility measured warn # 2 blocking. run not started.
Built for air-gapped networks, but not limited to them. Same deliverable either way.
Offline install from mirrored images, training on isolated hardware, delivery on physical media. No licence server, no telemetry, no callback anywhere in the path. This is the constraint the company was built around, and your firewall rules are the control, not our promises.
The run executes on the cluster in your data center or your VPC, against storage you control, with our engineers driving it. The right answer when residency rules bind the data but the network is not actually sealed.
Not everyone who needs a domain model is under supervisory rules. If nothing about your problem requires an air gap, we do not charge you for one: you bring the domain, we bring the compute and the training stack, and you still own the weights outright at the end.
Pretrained from scratch on our own GPUs. Every model ships with evals and a written account of where it fails.
Grounded financial analysis. Reads live market data from context and quotes it exactly, or tells you it does not have the number. Pretrained on SEC EDGAR filings and financial news. 9B parameters with 2.6B active, Apache-2.0, commercial use permitted.
weights, evals and limits →Search a codebase by describing what the code does, in plain English. Pretrained from scratch on 3.97B tokens of permissively licensed code, with every licence decision and every redaction written down. 110M parameters, runs on a CPU, Apache‑2.0.
weights, evals and limits →Analyzes COBOL programs to identify files, copybooks, data fields, called programs, and dead code. Trained on real and generated COBOL from The Stack. 1.5B parameters, Apache-2.0, commercial use permitted.
weights, evals and limits →What that looks like in use: an analyst drops a quarterly filing into the context window and asks for the segment margin. Kautilyaa returns the figure and the line it came from. Ask it for something the filing does not contain and it says it does not have that number. Across 150 test cases it produced no invented figures at all, which is the behaviour that decides whether a model is allowed anywhere near a client report.
On Kautilyaa we traced a silent optimizer bug in the training stack that had been quietly degrading every run: weight decay applied to normalisation gains, invisible in the loss curve, measurable only in the weights afterwards. Catching that class of failure before your compute is spent is what a Foundry engagement is actually for.
The frontier labs won't train inside your network or hand over weights. We do both.
Delivered as safetensors and GGUF with the eval harness and the configs to reproduce the run. There is no per-seat AI tax and no renewal date, so nobody gets to decide next year what you are allowed to keep running.
Corpus, configuration, and tokenizer checks run before a GPU hour is billed. When the run finishes we open the weights up and look at what is actually in them, then hand you the harness so you can repeat every check yourself.
Local inference by default, offline licence validation, your own SMTP and object storage. Nothing in the install path requires a call back to us. Block us at the perimeter and the software carries on working exactly as before.
A few hundred real examples expanded into training volume, or a corpus we curate for you from public sources. Starting with no dataset changes what the build costs and how long it takes. Plenty of teams start there.
We build in the 9B class, sparse where it helps, so the result serves on hardware you already own. There is no point delivering something that will not fit on the GPUs sitting in your rack.
Roughly half of our corpus audits conclude that retrieval or a fine-tune is enough, and we put that in writing. We are not charging for the week at present, so telling you to walk away costs us nothing.
If you can't send your data to a frontier lab, building a model is the only path left.
Classified and controlled-unclassified networks where air-gap is a hard requirement and offline delivery is the only install path.
Core banking, transaction, and PII data under supervisory rules that forbid third-party model processing.
Patient records and trial data, where auditors care about where the data physically sits rather than how well it is encrypted in transit.
Energy, utilities, and industrial operators running OT networks that are isolated from the internet by design.
If you have a pursuit with an on-premise or sovereign model requirement, we build it under your delivery. Your client ends up owning the weights outright. There is no telemetry and no callback in anything we hand over, so it should clear your security review without argument, and the audit report is written so it can go into your bid as it stands.
We read your corpus, train small models on it, and tell you what a build would cost. Half end with us recommending you don't build. Free for now.
Your domain, what data you actually have, where the build is allowed to run, and a scaling ladder on your corpus to see what size is justified.
The findings, a token budget, a recommended rung, and the eval targets a build would have to hit. Including the recommendation not to build, when that is the honest one.
On our GPUs, on yours, or inside your air gap. The eval targets are agreed in writing before the run starts, and if the delivered weights miss them the remediation work is ours, not billed to you.
A Foundry engagement is complete when the weights are in your hands. Everything below is a separate product that exists because we needed it ourselves to build models. Three of them are open source, so if you are weighing up a regulated build with a small vendor, you can go and read the code today.
If you also want the pipelines around the model: an agentic data platform that serves your weights on your own Ollama or vLLM and ships pipeline code to your Spark and Kafka. Self-hosted, air-gap capable, from $60,000 a year. Decline it and your model works exactly the same.
Parallel agent fleets and persistent memory for coding and data engineering, in your terminal. Free forever, bring your own keys.
A memory database for agents. Intent goes in, the right context comes out, packed to a token budget.
A typed, compiled language where tables, streams, and models are native types.