Owned AI infrastructure · built · under evaluation

Owned inference is an operating choice, not a hardware hobby.

Repurposed datacenter accelerators, a fifth card that can serve either host, one routing gateway, and a deliberate boundary between work that may leave and work that should remain under local control.

Why own it

Continuity

The model you built on is not yours

A frontier model I had in production was withdrawn on the vendor's schedule, not mine. Work that had been reproducible stopped being reproducible overnight, and the replacement cost more and behaved differently. That is not a defect in the vendor's business — it is the vendor's business. Weights on a disk I own cannot be deprecated out from under a client engagement.

Open the continuity record →
Economics

A metered bill against a fixed cost

Per-token pricing turns every experiment into a decision about whether to run it, and the industry's own pricing has moved sharply and without much notice. Hardware that is already paid for inverts that: the marginal cost of another hundred evaluation runs is electricity. That is the difference between testing what you can afford and testing what you should.

Open the tier architecture →
Control

What leaves is a decision, not a default

I am not trying to stop data leaving. I am trying to be the one who decides when it does. Plenty of my work should go out — a specification is sharper after two vendors have argued over it. Plenty should not: a client's operational schema may have no business leaving infrastructure the client controls. A local tier does not guarantee good data handling by itself; it makes storage, logging, retention, and network access policies we can inspect and enforce. Owning the tier underneath turns disclosure into a routing choice instead of a default.

Open the disclosure argument →

Local and commercial, on purpose

The farm is not there to end my use of frontier models. It is there so that using them is a choice with an alternative behind it. Both tiers sit behind the same gateway and both earn their traffic.

What the commercial tiers are for

Independent challenge is more useful than another confident answer

The frontier tier buys something owned hardware cannot cheaply produce: genuinely independent opinions. When a specification or a review finding matters, I put it in front of more than one commercial model and treat the disagreements as a reason to inspect the evidence. Multiple reviewers can still share the same blind spot, so agreement raises confidence only when the underlying claim can also be checked against source, tests, or the running system.

That is a harness technique, not a budget line. Cross-examination is the whole reason the commercial tiers stay wired in rather than being something to engineer away.

Open the quality-control loop →
How the built tier earns traffic

It earns a seat on measured work

The farm is built, and more powerful local models are now being tested on both servers. They are joining the panel, not replacing it by declaration. Each model still has to clear a recorded bar on the real work of its role before anything depends on it. Lower cost and local control are reasons to test a model; they are not reasons to trust one.

The next test is deliberately ordinary and difficult to fake: my daily programming. I will compare local results against the checks, reviews, and working software the job already produces. In parallel, the stronger local tier will help improve AI TraceVector Alpha itself — widening research and review without moving the acceptance decision away from evidence.

Open the model-selection decision →

Where it started

One card, one model, and a queue behind it.

The whole expansion traces to a single measured constraint. The original server had one 24 GB card. A 32B-class model quantised to fit takes about 18.5 GB on disk and peaks at 21–23 GB in service once you add a context window and a couple of decode slots. There is no room for a second one. Not "it would be slow" — it does not fit, so every agent in a multi-agent system waits its turn behind whichever model is currently resident. You cannot run a hierarchy of models on a machine that can only hold one of them at a time. Everything below is what it took to fix that without paying frontier prices to do it.

How it is put together

Three planes, borrowed from ordinary infrastructure design and applied to inference. The point of the split is that a consumer — the trading system, a coding harness, this site's own assistant — asks for a capability and never names a machine or a provider. Retiring a card, or moving a workload from local silicon to a metered API and back, is a routing change, not a rewrite.

What it actually cost

Buying cheap datacenter hardware means inheriting every assumption the people who designed it made about the datacenter it would live in. None of the following is in a build guide. Each one cost real days, and each one turns out to be the same lesson the software half of my work keeps teaching.

The two-week fault

A single bent pin in the CPU socket

The server board never completed POST. It parked at the same diagnostic code through two different CPUs, two thermal-paste jobs, a firmware upgrade and a BIOS upgrade. The management controller had been logging thermal-trip events, so the first two weeks went into cooling — and the cooling was fine. After the firmware update the failing boots logged nothing at all, and the controller showed the CPU temperature sensor and all eight memory temperature sensors as disabled. That was the real signal: the management sideband between the controller and the CPU runs through the socket, and it had gone silent. Thermal was the symptom. The disease was one flattened pin in the first row under the frame wall, drawn out into a gold ribbon lying across its neighbour.

An error that reports a plausible cause will absorb every hour you give it. The sensors that went quiet were worth more than the alarm that kept firing — absence is evidence, and almost nothing is built to alert on it.

The photograph that lied

Raked light, or you will not see it

Photographed from near-vertical, the damage vanished. A flattened pin presents no tip to the camera, so it renders as a slightly dark segment of an otherwise bright row — which reads as "nothing there." It only glints under raked light, and it took two independent oblique angles to establish that the anomaly was real and not an artefact. A canned-air puff from a distance ruled out the other candidate: a stray fibre would have blown away. It did not, so it was metal.

Design the test so that a negative result means something. "I looked and didn't see it" is not a measurement if the method cannot resolve the defect.

The card killer

Two 8-pin connectors that mate and must not

These accelerators take an EPS (CPU-style) 8-pin. Graphics cards normally take a PCIe 8-pin. Same appearance, physically matable, different wiring — four 12 V pins versus three, in different positions. Feed one the wrong way and 12 V lands on the card's ground pins: instant death, not covered by warranty. The keying is moulded into the card and cannot be wrong, but a no-brand adapter can have perfect keying with wires crimped into the wrong cavities. It mates, it photographs correctly, and it kills the card.

My first instinct was to check that the adapter ran each pin straight through to the same position. That is exactly backwards — a straight-through cable IS a PCIe cable. The adapter's whole job is to move the conductors. The thing to verify is never "do the ends match," it is "does the voltage land where the load expects it."

The meter

Silence has to mean something

The test is continuity from a known-good ground to each of the eight pins, using the power supply's bare chassis and the card's own bracket as references — count the beeps. Four means EPS and correct; five means PCIe wiring and stop. The step people skip is the self-test: touch the probes together first and confirm the meter beeps. A dead battery gives silence on every pin, and silence reads exactly like "no connection."

Any instrument whose failure mode looks like a passing result has to be proven before it is trusted. This is the multimeter version of a test suite that passes because it never ran.

The power supply

Pigtails, and cables that mate across brands

Modular PSU cables are pigtailed — one lead ends in two connectors. Using both heads of a single cable to feed one 250 W card pushes that current through copper rated for 150 W. Correct is one head from each of two separate cables. Worse: two manufacturers' modular cables physically mate at the PSU end while carrying entirely different pinouts, with no key to stop you. Nothing about the fit warns you. And the cables that shipped in the box with the cards were server-riser cables, not ATX-safe, despite being labelled for a GPU.

A connector that fits is not a connector that is correct. Every place this build could destroy hardware, it did so through a joint that mated perfectly.

The physical constraint

The card does not fit — and that made it movable

On the workstation, none of the three slots could take a second full-size card: the only CPU-attached x16 was occupied, the middle slot turned out to be a closed-ended x1 that a x16 edge physically cannot enter, and the bottom slot was long enough but collided with the power-supply shroud. So the fifth V100 lives outside both cases, on an OCuLink dock over a x4 link. The cable can place that 32 GB card beside the four-V100 server when a job needs a 160 GB pool, or beside the 24 GB RTX 3090 workstation when a smaller host needs 56 GB. The link is a quarter of the width — and it barely matters, because splitting a model across cards passes only small activations between them. The one real cost is loading weights off disk, once per load.

Measure what the bottleneck actually is. A specification that looks four times worse can be irrelevant to your workload, and finding that out is what lets you use hardware other people have ruled out.

And one from the software side

Renaming a directory broke the inference server.

Binaries built in a build tree bake an absolute library path into themselves. A rebuild for the second accelerator generation went into a new directory, the old one was moved aside, the new one renamed into its place — and every model spawn failed, because the path inside the binary no longer pointed at anything. The fix is to never rename a built directory: build beside it, verify, then repoint a symlink. Which is the same rule as the hardware ones — the artefact carries assumptions you did not write down, and moving it does not move them.

Read the incident →

Bigger was not better. Measured was.

Old hardware only works as a strategy if you know which model genuinely earns each role, so models here are assigned by a recorded bake-off rather than by parameter count. Two results from that record shaped the buying decisions:

That is the whole argument for the farm in one line: if the model that wins a role is chosen by measurement rather than by size, then owned hardware running the right mid-sized model is not a compromise against a frontier API. It is the better answer, and it keeps running when somebody else's roadmap changes.

Honest status, August 2026. The farm is built. Both servers are running local models, and more powerful candidates are being tested on each. The fifth V100 can be connected to the 128 GB V100 server or the 24 GB RTX 3090 workstation according to the job. The next evidence comes from daily programming and from using the local tier to improve AI TraceVector Alpha. I do not have final comparative benchmarks yet; when I do, this page gets the disappointing results as well as the wins.

Thinking about running your own?

If the reason is confidentiality, cost predictability, or not wanting a vendor's roadmap inside your product — those are good reasons, and the path is more navigable than the failure stories above make it sound. I have already paid for the map.