Quaestor, our deep research engine, was given a long-running, challenging development project: close an internal need, and prove itself on a real-world task. This is the full record of how it did it.
Quaestor researched the subject from primary sources, developed a strategy per subject lane, defined and deployed multiple coordinators, agent swarms and verifiers, defended its own findings through adversarial review. It then designed and tuned the result under measurement, and, through continuous improvement, designed and trained its own replacement workers while the campaign was still running.
We shipped a consumer-focused version of the resulting engine as GInfer, open source under Apache 2.0.
A worked example of the Quaestor deep research engine, applied to a subject of our own: designing and building the inference systems we run internally. A consumer-focused build of the result was released to the open source and local hosting community as GInfer.
We needed a production inference engine for our own agentic workloads, and the work involved was not confined to kernels. It covered the runtime structure, the K/V system, per-SM hot paths aligned to what each target GPU supports, and the quantization and calibration of every target model against that same hardware profile.
Work of that shape is normally distributed across a team of specialists over several quarters. Each decision depends on documentation nobody has read in full, on hardware behavior nobody has measured in full, and on tradeoffs that only become visible once something is running. We had a second reason for choosing it: we wanted a subject where the answer could be checked.
Quaestor begins by decomposing the subject into the questions that would settle it, then dispatching swarms to answer them from primary sources — hardware documentation, model specifications, CUDA and kernel-tuning references, and the published literature on serving and K/V design.
Findings are challenged before they are accepted. Standing review swarms test correctness, fact, and completeness, and can retrieve either a cited summary or the full document from the evidence locker. Unresolved questions return to research with a narrower scope.
Design work starts only after the Overseer accepts the research package. Candidates then carry a falsifiable hypothesis and the measurement that would disprove it, and every candidate leaves an evidence packet behind whether it is accepted, deferred, or rejected.
Quaestor improves through its loops. Each accepted win becomes the next baseline, a fresh candidate set is identified against it, and the cycle repeats until a full sweep finds nothing worth pursuing.
The models doing the work improve alongside it. Timings, accepted candidates, and labeled rejections stream to the Dataset Coordinator throughout the campaign and become the training corpus for the next Quaestor Expert Model — a bespoke worker, trained by Quaestor on its own output.
Each generation is measurably better at two specific tasks: identifying candidates worth a lane, and tuning them once found. Fewer wasted lanes and shorter paths to an accepted win compound across the remaining campaign, and the expert model is retained as an artifact of the engagement.
Coverage target
Parallel, long-running
Full documents
Cited summaries
Score, clean, assemble
Built continuously as the lanes run
Bespoke, domain-trained — v1 → vN
sm_80 · 86 · 89 · 120a ×2
Never shared across targets
Lane constants
Engine · K/V · quantization · base kernels — run twice
Test, iterate, decide
Deep timings and instrumentation
Ranked from timings and instrumentation
Own eval
Own eval
Own eval
All three lanes, one rejection-first gate stack
A bespoke model, custom designed and trained by Quaestor itself to perform one defined set of tasks. This campaign produced an expert in whole-engine design — runtime architecture, per-SM hot paths, K/V system design, target quantization and calibration, and kernel tuning. Quaestor continuously self-trains the worker model on its own targeted outputs. Each iteration the model self-improves in its function and is preserved as a set of trained weights as an artifact of the campaign.
Every lane host running the inference engine, the K/V system, and the tuned kernels
From intake decomposition
Retrieves, reads, and extracts
Every assertion names its source
Published in-project
The claim is attacked before it is accepted
Is the claim true as stated?
Does the source actually say it?
What has been left out?
Built into the harness. A claim does not advance on assertion
Accepted · corrected · unproven
The reviewer may pull the cited summary, or the full document from the evidence locker
The claim is rewritten to match the evidence
The question re-enters the swarm with a narrower scope
The researcher defends the claim as written
Stage timers, occupancy, residency, power, host-side stalls
What was proposed, and the measurement that would disprove it
The actual change, bound to the candidate that made it
Which gate passed, which failed, and in what order
Build · correctness · route · performance · resource
What the harness challenged, and how it was resolved
The evidence each decision was grounded in
Signal quality and novelty per packet
Normalize, validate, deduplicate
Outcome class and reason per candidate
Into the working set, with provenance intact
Wins against rejections, so the model learns both
Curated corpus, reviewed by the harness
Fine-tune the next expert model on the campaign’s own output
Held-out campaign work, judged by the same review harness
The new expert takes the lanes; the previous one is preserved
A worker that has seen ten thousand labeled rejections stops proposing candidates that resemble them. Lanes are finite, so every lane not wasted is a lane spent on something that might win.
The gain is cumulative and measured in the same currency as the work: accepted wins per lane, and lanes per accepted win.
Deep analysis is expensive because most of the exploration is wasted. Lower the waste and the same budget buys a longer chain of reasoning: more questions pursued, more candidates tested, more of the subject actually covered.
This is what enables Quaestor to shorten a campaign by weeks and still produce a defensible result that would take a team months to accomplish.
Each campaign leaves two things: the result, and a bespoke model that is measurably better at producing that class of result. The second is preserved as weights and carried into the next campaign on the same subject.
Future research on a similar subject then benefits from the experience of the last.
Owns one GPU × model × quantization, and nothing else
The standing best, replay-verified
The patch, the result, the decision, and a label for why it failed. For every worker, accepted or not.
Each worker has its own worktree, build root, and profile root, one coherent mechanism, and a falsifiable hypothesis naming the measurement it targets.
Independent, and before any timing
A/B or AB/BA against the baseline
An unexplained stall triggers a profiler sentinel
Correctness, latency, throughput, prefix routes, rank parity
Becomes the next baseline. Unfinished siblings are marked stale and requalify.
Positive but under threshold: to the residual pool, rechecked against later baselines.
Labeled by outcome class and kept in the training dataset.
The DFlash2 module as its own subject, not a sub-task of kernel tuning
What a draft model must do at each concurrency and context length
Draft acceptance, window sizing, compute-type support per SM
Swarms against hardware docs, model specs, and the speculative-decoding literature
Every citation retrievable as summary or full document, at any gate
Adversarially reviewed for correctness, fact, and completeness
One DFlash2 variant per SM, quantized to that SM's supported compute type rather than to a single format applied everywhere.
Each variant carries kernels tuned for it specifically: the same candidate, gate, and acceptance machinery the main kernel path uses, run against the draft model instead of the target.
Two changes came out of the research loop, both aimed at the concurrency and context profile agentic work actually produces.
The draft window is sized to the live concurrency and context length rather than fixed at build time.
Single-stream stays sequential. Above C1 the drafts run in parallel, which is where sustained agentic serving spends its time.
Objective, goals, and conditions
Ordered work, each with its acceptance criteria
To the coordinator that owns the phase
Scope confirmed before any swarm is built
Progress, cost, and coverage while it runs
The artifact, with its review record attached
The Overseer alone accepts. Nothing else can promote a package.
An independent numerical oracle runs before any timing, so an incorrect candidate is discarded before it is profiled. Matched timing, a utilization sweep, and the full matrix follow in ascending order of cost.
Most candidates are eliminated cheaply, which is what allows a finite exploration budget to cover a large candidate space.
Each candidate changes exactly one thing and runs in its own worktree with its own build and profile roots. Nothing is shared with a sibling.
Every measurement can be traced to a single cause. A result nobody can attribute is a result nobody can act on, and the harness does not produce them.
Findings are filed with citations. Review can pull a cited summary or the full source document at any gate, and a finding that cannot be traced back is returned on the first pass.
The output survives a hostile read. For a regulated or contested decision, that property is the deliverable.
Progress is measured against the question tree the subject was decomposed into. The loop ends when the tree is covered and a full sweep returns nothing worth pursuing.
The engagement finishes on a finding rather than on a date, and the scope of what was examined is stated rather than implied.
Timings, accepted candidates, and labeled rejections stream to curation throughout the campaign and train the next generation of workers mid-run.
Each generation wastes fewer lanes and reaches an accepted result in fewer steps. Long-horizon work becomes affordable because the horizon gets cheaper as it extends.
Only the Overseer can promote a package, and it holds no other authority. No coordinator has an alternative route forward, whatever the schedule pressure.
Rigor does not depend on discipline holding under deadline. It is the only available path.
C1 at 128K context
C1 at 128K context
Decode aggregate, against the legacy engine
Decode per request, tok/s. Twice the concurrency, and each request still completes faster.
| Metric | Legacy low | Post-tuning high | Improvement |
|---|---|---|---|
| Prefill · PP | 614.214 tok/s | 1,165.255 tok/s | 1.90× · +89.7% |
| Decode | 106.743 tok/s | 271.109 tok/s | 2.54× · +154.0% |
| Concurrency | Prefill · legacy → tuned · tok/s | Gain | Decode aggregate · tok/s | Decode per request · tok/s | Gain |
|---|---|---|---|---|---|
| C1 | 3,035.944 → 3,453.983 | +13.8% | 154.450 → 172.476 | same as aggregate | +11.7% |
| C4 | 3,022.105 → 3,433.804 | +13.6% | 195.694 → 317.628 | 48.924 → 79.407 | +62.3% |
| C8 | 3,017.146 → 3,394.819 | +12.5% | 287.486 → 441.372 | 35.936 → 55.172 | +53.5% |
The tuned engine on every target it was built for: two models, six GPU configurations from SM80 to Blackwell, at one, four and eight requests in flight. The legacy comparison above shows how far the campaign moved one card; this is where the result stands on all of them.
Qwen 3.8 27B · RTX PRO 6000 ×2 · C8
Qwen 3.8 27B · RTX PRO 6000 · C8
Muse Glimmer 30B · RTX PRO 6000 · C1
Muse Glimmer 30B · RTX PRO 6000 ×2 · C1
| GPU | SM | TP | C | Cold PP t/s | PP t/s | TG t/s | TTFT ms |
|---|---|---|---|---|---|---|---|
| CMP 170HX | 80 | 1 | C1 | 1,678 | 774,281 | 126.2 | 2.65 |
| C4 | 1,673 | 1,141,910 | 345.7 | 7.17 | |||
| C8 | 1,672 | 1,244,242 | 521.2 | 13.17 | |||
| RTX 3090 | 86 | 1 | C1 | 857 | 628,060 | 149.0 | 3.26 |
| C4 | 863 | 1,071,410 | 434.3 | 7.65 | |||
| C8 | 858 | 1,101,520 | 614.5 | 14.87 | |||
| RTX 4090 | 89 | 1 | C1 | 2,007 | 683,153 | 321.0 | 3.00 |
| C4 | 2,004 | 1,132,717 | 686.8 | 7.23 | |||
| C8 | 2,003 | 937,059 | 1,127.9 | 17.48 | |||
| RTX 5090 | 120 | 1 | C1 | 8,204 | 796,640 | 654.6 | 2.57 |
| C4 | 8,197 | 1,052,611 | 1,596.6 | 7.78 | |||
| C8 | 8,179 | 1,044,175 | 2,143.7 | 15.69 | |||
| RTX PRO 6000 | 120 | 1 | C1 | 10,117 | 915,598 | 657.7 | 2.24 |
| C4 | 10,090 | 1,709,351 | 1,735.7 | 4.79 | |||
| C8 | 10,051 | 1,560,745 | 2,483.2 | 10.50 | |||
| RTX PRO 6000 ×2 | 120 | 2 | C1 | 8,253 | 689,084 | 951.1 | 2.97 |
| C4 | 8,253 | 1,448,022 | 2,443.1 | 5.66 | |||
| C8 | 8,240 | 1,398,760 | 3,224.3 | 11.71 |
| GPU | SM | TP | C | Cold PP t/s | PP t/s | TG t/s | TTFT ms |
|---|---|---|---|---|---|---|---|
| CMP 170HX | 80 | 1 | C1 | 3,571 | 716,054 | 132.4 | 2.86 |
| C4 | 3,596 | 1,385,127 | 287.2 | 5.91 | |||
| C8 | 3,564 | 1,048,915 | 422.1 | 15.62 | |||
| RTX 3090 | 86 | 1 | C1 | 2,250 | 688,968 | 190.2 | 2.97 |
| C4 | 2,233 | 1,406,472 | 371.4 | 5.82 | |||
| C8 | 2,212 | 1,285,274 | 481.0 | 12.75 | |||
| RTX 4090 | 89 | 1 | C1 | 5,578 | 690,922 | 360.2 | 2.96 |
| C4 | 5,575 | 1,219,020 | 767.8 | 6.72 | |||
| C8 | 5,571 | 1,401,778 | 1,291.9 | 11.69 | |||
| RTX 5090 | 120 | 1 | C1 | 28,761 | 680,102 | 559.0 | 3.01 |
| C4 | 28,543 | 924,663 | 1,382.1 | 8.86 | |||
| C8 | 28,386 | 850,446 | 2,303.3 | 19.27 | |||
| RTX PRO 6000 | 120 | 1 | C1 | 34,514 | 890,375 | 579.4 | 2.30 |
| C4 | 34,576 | 1,583,770 | 1,442.1 | 5.17 | |||
| C8 | 34,442 | 1,322,752 | 2,431.3 | 12.39 | |||
| RTX PRO 6000 ×2 | 120 | 2 | C1 | 17,983 | 1,024,483 | 767.0 | 2.00 |
| C4 | 17,955 | 1,567,132 | 1,813.2 | 5.23 | |||
| C8 | 17,980 | 1,514,363 | 2,093.7 | 10.82 |
The campaign produced a complete inference engine: runtime structure, the K/V system, per-SM hot paths aligned to each target's capability, per-SM quantization and calibration of the target models, the DFlash2 speculative path, and the tuned kernels beneath all of it.
It also produced a Quaestor Expert Model: a bespoke worker trained by the engine on its own output, preserved as weights and carried forward as an artifact of the engagement.
GInfer is measured against replay-verified baselines and frozen per-SM artifacts, under the workload profile agentic serving produces: high concurrency, long shared prefixes, repeated tool turns, and sustained high K/V occupancy. At 128K context and concurrency one, prefill reached 1.90× and decode 2.54× against the legacy engine. Under concurrency, decode aggregate gained 62.3% at C4 and 53.5% at C8.
These figures are internal measurement of GInfer, pending independent verification.
One campaign, on one subject, does not establish that the method generalizes. It establishes that the machinery runs end to end on a hard, measurable problem and produces an auditable result, which is the narrower claim we are making.
The representation and architecture behind Tessera v1 remain undisclosed pending our initial patent filings. Nothing in this case study depends on them.
The mandate was whole-engine: runtime architecture, per-SM hot paths matched to each SM’s hardware capability, the K/V system, and target quantization chosen for compute capability. Kernels came last.
Each pass preserves a set of trained weights. The expertise is retained rather than recomputed, and the next campaign in the domain starts from it.
Frozen target models, isolated workers, one mechanism each, and cheap gates before expensive ones. A measured improvement belongs to a named change.
Rejections carry as much training signal as wins, which is why the worker model improved while the kernels did.
Quaestor is available for custom research applications, scoped per engagement or via annual licenses — whether you require diligence, regulatory impact, contract and obligation review, portfolio assessment, M&A assessment, IT landscape assessment, complex migration planning, complex codebase reviews, compliance support, or any other domain where the analysis is broad, sourced, and will be contested. Quaestor will reduce your cost and produce your deliverables faster, with entirely defensible results.
A worked example of the Quaestor deep research engine, applied to a subject of our own: designing and building the inference systems we run internally. A consumer-focused build of the result was released to the open source and local hosting community as GInfer.
We needed a production inference engine for our own agentic workloads, and the work involved was not confined to kernels. It covered the runtime structure, the K/V system, per-SM hot paths aligned to what each target GPU supports, and the quantization and calibration of every target model against that same hardware profile.
Work of that shape is normally distributed across a team of specialists over several quarters. Each decision depends on documentation nobody has read in full, on hardware behavior nobody has measured in full, and on tradeoffs that only become visible once something is running. We had a second reason for choosing it: we wanted a subject where the answer could be checked.
Quaestor begins by decomposing the subject into the questions that would settle it, then dispatching swarms to answer them from primary sources — hardware documentation, model specifications, CUDA and kernel-tuning references, and the published literature on serving and K/V design.
Findings are challenged before they are accepted. Standing review swarms test correctness, fact, and completeness, and can retrieve either a cited summary or the full document from the evidence locker. Unresolved questions return to research with a narrower scope.
Design work starts only after the Overseer accepts the research package. Candidates then carry a falsifiable hypothesis and the measurement that would disprove it, and every candidate leaves an evidence packet behind whether it is accepted, deferred, or rejected.
Quaestor improves through its loops. Each accepted win becomes the next baseline, a fresh candidate set is identified against it, and the cycle repeats until a full sweep finds nothing worth pursuing.
The models doing the work improve alongside it. Timings, accepted candidates, and labeled rejections stream to the Dataset Coordinator throughout the campaign and become the training corpus for the next Quaestor Expert Model — a bespoke worker, trained by Quaestor on its own output.
Each generation is measurably better at two specific tasks: identifying candidates worth a lane, and tuning them once found. Fewer wasted lanes and shorter paths to an accepted win compound across the remaining campaign, and the expert model is retained as an artifact of the engagement.