sectilelabs
Case study · Quaestor

We used our universal deep research engine to solve a technically challenging and complex issue that its expert models had never seen before: handing off the end-to-end development of our own internal inference engine.

Quaestor, our deep research engine, was given a long-running, challenging development project: close an internal need, and prove itself on a real-world task. This is the full record of how it did it.

About Quaestor
In short

Quaestor researched the subject from primary sources, developed a strategy per subject lane, defined and deployed multiple coordinators, agent swarms and verifiers, defended its own findings through adversarial review. It then designed and tuned the result under measurement, and, through continuous improvement, designed and trained its own replacement workers while the campaign was still running.

We shipped a consumer-focused version of the resulting engine as GInfer, open source under Apache 2.0.

The subject
A complete inference engine, tuned per GPU
Quaestor built the new runtime structures, a state-of-the-art K/V system, per-SM hot paths, and per-SM quantization of every target model in a complex deliverable matrix.
The method
Research, design, implement, improve
The subject was decomposed into questions and answered from primary sources under adversarial review. Design began only on accepted research, every change was built and measured against a frozen baseline, and each accepted win became the next baseline.
The difference
Models that train themselves on the work they do
Every timing, and every candidate, successful or not, was labeled and entered the dataset pipeline as training data mid-campaign, so each generation is more efficient than the last: better results, a faster pace, and lower runtime and costs.
01
Quaestor · GInfer Development

Ginfer development, a Quaestor practical application case study

A worked example of the Quaestor deep research engine, applied to a subject of our own: designing and building the inference systems we run internally. A consumer-focused build of the result was released to the open source and local hosting community as GInfer.

4 weeksCampaign duration
6Phases under one control bus
857Questions from decomposition
204Source documents gathered at intake
1,632 / 497Candidates tested / accepted
4Expert model generations
01 · The mandate

We needed a production inference engine for our own agentic workloads, and the work involved was not confined to kernels. It covered the runtime structure, the K/V system, per-SM hot paths aligned to what each target GPU supports, and the quantization and calibration of every target model against that same hardware profile.

Work of that shape is normally distributed across a team of specialists over several quarters. Each decision depends on documentation nobody has read in full, on hardware behavior nobody has measured in full, and on tradeoffs that only become visible once something is running. We had a second reason for choosing it: we wanted a subject where the answer could be checked.

02 · Research as a development method

Quaestor begins by decomposing the subject into the questions that would settle it, then dispatching swarms to answer them from primary sources — hardware documentation, model specifications, CUDA and kernel-tuning references, and the published literature on serving and K/V design.

Findings are challenged before they are accepted. Standing review swarms test correctness, fact, and completeness, and can retrieve either a cited summary or the full document from the evidence locker. Unresolved questions return to research with a narrower scope.

Design work starts only after the Overseer accepts the research package. Candidates then carry a falsifiable hypothesis and the measurement that would disprove it, and every candidate leaves an evidence packet behind whether it is accepted, deferred, or rejected.

03 · Why the cost curve falls

Quaestor improves through its loops. Each accepted win becomes the next baseline, a fresh candidate set is identified against it, and the cycle repeats until a full sweep finds nothing worth pursuing.

The models doing the work improve alongside it. Timings, accepted candidates, and labeled rejections stream to the Dataset Coordinator throughout the campaign and become the training corpus for the next Quaestor Expert Model — a bespoke worker, trained by Quaestor on its own output.

Each generation is measurably better at two specific tasks: identifying candidates worth a lane, and tuning them once found. Fewer wasted lanes and shorter paths to an accepted win compound across the remaining campaign, and the expert model is retained as an artifact of the engagement.

Quaestor Overseer Model: internally built and trained model specifically to run Quaestor campaigns
Control bus
01 · Intake
Subject

Design, build, and optimize a complete inference engine for every target GPU

Goal decomposition

Question set

Coverage target

Unresolved — back to decomposition
The mandate

Design the engine, and custom kernels

  • Engine architecture and runtime structure
  • Per-SM hot paths, aligned to what each SM can actually do
  • The K/V system — representation, segments, and routes
  • Per-SM quantization of the target models, chosen for compute capability
  • Calibration bound to each target, never shared across them
  • Kernels, tuned last and measured against a frozen artifact
02 · Knowledge
Coordinator

Research coordinator

Research agent swarm

Parallel, long-running

Evidence locker

Full documents

Research dossier

Cited summaries

Coverage gap re-enters the swarm
03 · Expert model factory
Coordinator

Dataset coordinator

Curation swarm

Score, clean, assemble

Dataset v1 → vN

Built continuously as the lanes run

Coordinator

Training coordinator

Validation loop

Quaestor Expert Model

Bespoke, domain-trained — v1 → vN

Fails validation — back to training
04 · Target prep
Coordinator

Target coordinator

Quantize per SM

sm_80 · 86 · 89 · 120a ×2

Calibrate per target

Never shared across targets

Frozen per-SM artifacts

Lane constants

Parallel campaign path

DFlash2 speculative decoding

  • Quantized per SM to that SM’s supported compute type
  • Its own per-SM tuned kernels
  • Sliding windows sized to concurrency and context
  • Parallel draft runs above C1
DFlash2 feeds generation
05 · Design
Coordinator

Campaign model

Design loops

Engine · K/V · quantization · base kernels — run twice

Construction loops

Test, iterate, decide

Accepted design

Change engine / K/V / requantize
06 · Tuning and improvement
Coordinator

Tuning coordinator

Baseline testing

Deep timings and instrumentation

Candidate identification

Ranked from timings and instrumentation

Kernel

Per SM × model

Own eval

Engine

Per-SM patch

Own eval

K/V

Per-SM tuning

Own eval

Merged candidate evaluation

All three lanes, one rejection-first gate stack

Accept · defer · reject

New baseline

Residual pool

Evidence packets

All three dispatched in parallel · new baseline advances · residual rechecked
Feedback and self-improvement lanesEvidence is collected continuously, accepted and rejected alike
Continuous self-training worker models

Quaestor Expert Model

A bespoke model, custom designed and trained by Quaestor itself to perform one defined set of tasks. This campaign produced an expert in whole-engine design — runtime architecture, per-SM hot paths, K/V system design, target quantization and calibration, and kernel tuning. Quaestor continuously self-trains the worker model on its own targeted outputs. Each iteration the model self-improves in its function and is preserved as a set of trained weights as an artifact of the campaign.

Every lane host

Worker instances

Every lane host running the inference engine, the K/V system, and the tuned kernels

Review loop harness
Fixed process, standing team — correctness, fact-checking, completeness, with scoring gates built in
Evidence ladder
Cited summary, or the full document from the locker
Resolution
Correct it, find more, or justify the position
Expert model vN drives every lane · new expert accepted → re-sweep
Harness reads the evidence locker at any gate
Every accepted expert model re-sweeps design and tuning. The campaign ends when a sweep finds nothing left to improve.
Termination
Gray fill coordinatorWhite processTeal left edge storeInk persistent service or modelSolid teal dispatch and flowDashed gray return or retry
02Research servicesEvery finding is challenged before it is accepted. Review maintains its own standing swarms for correctness, fact-checking, and completeness, with the authority to retrieve the underlying document at any gate.
One question
Input

Open question

From intake decomposition

Swarm

Research agent

Retrieves, reads, and extracts

Output

Claim with citation

Every assertion names its source

Store

Dossier entry

Published in-project

Corrections return to the swarm until the gate closes
Standing team
Standing team

Adversarial review

The claim is attacked before it is accepted

Mini-swarm

Correctness

Is the claim true as stated?

Mini-swarm

Fact-checking

Does the source actually say it?

Mini-swarm

Completeness

What has been left out?

Fixed process

Scoring gates

Built into the harness. A claim does not advance on assertion

Verdict

Accepted · corrected · unproven

Verdict routes to one of three
Resolution
Available at every gate

Evidence ladder

The reviewer may pull the cited summary, or the full document from the evidence locker

01

Correct it

The claim is rewritten to match the evidence

02

Find additional documentation

The question re-enters the swarm with a narrower scope

03

Justify the position

The researcher defends the claim as written

03Dataset generation and trainingQuaestor harvests its own training data from work it has already done, then trains the next generation of its own workers on it while the campaign is still running. This is the mechanism behind its economics, and the clearest difference between Quaestor and an agent that forgets.
HarvestProduced by work the campaign was already doing, never a separate collection effort
01

Timings

Stage timers, occupancy, residency, power, host-side stalls

02

Hypotheses

What was proposed, and the measurement that would disprove it

03

Patches

The actual change, bound to the candidate that made it

04

Gate outcomes

Which gate passed, which failed, and in what order

05

Rejection labels

Build · correctness · route · performance · resource

06

Review objections

What the harness challenged, and how it was resolved

07

Citations

The evidence each decision was grounded in

Continuous evidence stream — written as the lanes run, never batched at the end
Campaign-long
CurateRuns during the campaign, not after it
01

Score

Signal quality and novelty per packet

02

Clean

Normalize, validate, deduplicate

03

Label

Outcome class and reason per candidate

04

Assemble

Into the working set, with provenance intact

05

Balance

Wins against rejections, so the model learns both

Store

Dataset vN

Curated corpus, reviewed by the harness

Train
07

Train

Fine-tune the next expert model on the campaign’s own output

08

Validate

Held-out campaign work, judged by the same review harness

09

Promote

The new expert takes the lanes; the previous one is preserved

Fails validation — the incumbent keeps working
The improved expert produces the next generation of evidence. The loop closes on itself
Efficiency

Each generation costs less than the last

A worker that has seen ten thousand labeled rejections stops proposing candidates that resemble them. Lanes are finite, so every lane not wasted is a lane spent on something that might win.

The gain is cumulative and measured in the same currency as the work: accepted wins per lane, and lanes per accepted win.

Depth

Long-running analysis becomes affordable

Deep analysis is expensive because most of the exploration is wasted. Lower the waste and the same budget buys a longer chain of reasoning: more questions pursued, more candidates tested, more of the subject actually covered.

This is what enables Quaestor to shorten a campaign by weeks and still produce a defensible result that would take a team months to accomplish.

The artifact

The expert model is a deliverable

Each campaign leaves two things: the result, and a bespoke model that is measurably better at producing that class of result. The second is preserved as weights and carried into the next campaign on the same subject.

Future research on a similar subject then benefits from the experience of the last.

04Worker loopsA single lane, in detail. Many workers explore the same baseline simultaneously, each in isolation, because isolation is what allows a measurement to be attributed to one mechanism.
One lane

Lane controller

Owns one GPU × model × quantization, and nothing else

Read-only

Baseline

The standing best, replay-verified

Baseline advances — siblings requalify
Always written

Evidence packet

The patch, the result, the decision, and a label for why it failed. For every worker, accepted or not.

WorkersIsolated

Each worker has its own worktree, build root, and profile root, one coherent mechanism, and a falsifiable hypothesis naming the measurement it targets.

Isolated

Worker 01

Isolated

Worker 02

Isolated

Worker 03

Isolated

Worker 04

Isolated

Worker 05

Isolated

Worker n

GatesCheap gates before expensive ones
01

Numerical oracle

Independent, and before any timing

02

Matched timing

A/B or AB/BA against the baseline

03

Utilization sweep

An unexplained stall triggers a profiler sentinel

04

Cycle-full matrix

Correctness, latency, throughput, prefix routes, rank parity

Disposition
01

Accept

Becomes the next baseline. Unfinished siblings are marked stale and requalify.

02

Defer

Positive but under threshold: to the residual pool, rechecked against later baselines.

03

Reject

Labeled by outcome class and kept in the training dataset.

05The DFlash2 pathSpeculative decoding ran as its own campaign path. The same machinery as the kernel path (subject, goal decomposition, question sets, research loop, evidence locker, adversarial review) applied to a different subject.
01

Subject

The DFlash2 module as its own subject, not a sub-task of kernel tuning

02

Goal decomposition

What a draft model must do at each concurrency and context length

03

Question sets

Draft acceptance, window sizing, compute-type support per SM

04

Research loop

Swarms against hardware docs, model specs, and the speculative-decoding literature

05

Evidence locker

Every citation retrievable as summary or full document, at any gate

06

Accepted research

Adversarially reviewed for correctness, fact, and completeness

Accepted research drives all three
Per target

Quantized to what the silicon supports

One DFlash2 variant per SM, quantized to that SM's supported compute type rather than to a single format applied everywhere.

INT4INT8FP8NVFP4BF16
SM80 — CMP 170HX, A100
SM86 — RTX 3090
SM89 — RTX 4090
SM120a — RTX 5090
SM120a — RTX PRO 6000
Per target

Its own tuned kernels

Each variant carries kernels tuned for it specifically: the same candidate, gate, and acceptance machinery the main kernel path uses, run against the draft model instead of the target.

Same gate order — numerical oracle before any timing, then matched A/B
Same isolation — one worktree, one build root, one mechanism per candidate
Same disposition — accept, defer to the residual pool, or reject with a labeled reason
Measured against — the frozen per-SM artifact, never a shared baseline
Runtime behavior

Sized to the workload, not to a default

Two changes came out of the research loop, both aimed at the concurrency and context profile agentic work actually produces.

Dynamic sliding windows

The draft window is sized to the live concurrency and context length rather than fixed at build time.

Parallel draft runs above C1

Single-stream stays sequential. Above C1 the drafts run in parallel, which is where sustained agentic serving spends its time.

Every DFlash2 decision tracks the same target as the rest of the campaign: a production engine for agentic serving, tuned for higher concurrency and longer context, not for single-stream benchmarks
GInfer goal
06Overseer dispatchEvery phase enters and leaves through the same control plane. Only the Overseer can promote a package, and that constraint is what makes review mandatory in practice.
One control plane for the whole campaign
01

Intake

Objective, goals, and conditions

02

Phase queue

Ordered work, each with its acceptance criteria

03

Dispatch

To the coordinator that owns the phase

04

Coordinator acknowledges

Scope confirmed before any swarm is built

05

Telemetry

Progress, cost, and coverage while it runs

06

Package presented

The artifact, with its review record attached

07

Accept authority

The Overseer alone accepts. Nothing else can promote a package.

Returned — dispatched again with the harness findings
Accepted — the next phase leaves the queue
07Why it worksSix mechanisms, consolidated. Each is enforced on every package by the same harness, which is what separates the method from a well-intentioned process document.
01 · Mechanism

Rejection-first gating

Cheap tests before expensive ones

An independent numerical oracle runs before any timing, so an incorrect candidate is discarded before it is profiled. Matched timing, a utilization sweep, and the full matrix follow in ascending order of cost.

Most candidates are eliminated cheaply, which is what allows a finite exploration budget to cover a large candidate space.

02 · Mechanism

Attributable measurement

One mechanism per number

Each candidate changes exactly one thing and runs in its own worktree with its own build and profile roots. Nothing is shared with a sibling.

Every measurement can be traced to a single cause. A result nobody can attribute is a result nobody can act on, and the harness does not produce them.

03 · Mechanism

A retrievable evidence ladder

Nothing advances on assertion

Findings are filed with citations. Review can pull a cited summary or the full source document at any gate, and a finding that cannot be traced back is returned on the first pass.

The output survives a hostile read. For a regulated or contested decision, that property is the deliverable.

04 · Mechanism

Coverage-driven termination

An exit condition, not a budget

Progress is measured against the question tree the subject was decomposed into. The loop ends when the tree is covered and a full sweep returns nothing worth pursuing.

The engagement finishes on a finding rather than on a date, and the scope of what was examined is stated rather than implied.

05 · Mechanism

Continuous self-training

Cost per result falls during the run

Timings, accepted candidates, and labeled rejections stream to curation throughout the campaign and train the next generation of workers mid-run.

Each generation wastes fewer lanes and reaches an accepted result in fewer steps. Long-horizon work becomes affordable because the horizon gets cheaper as it extends.

06 · Mechanism

Centralized acceptance

The harness cannot be bypassed

Only the Overseer can promote a package, and it holds no other authority. No coordinator has an alternative route forward, whatever the schedule pressure.

Rigor does not depend on discipline holding under deadline. It is the only available path.

Quaestor works because all six are enforced by the same harness on every package, so none of them is the first thing to be dropped when a campaign runs long.
Why it holds
08ResultsWhat the campaign produced, how it is being measured, and what it has not yet earned the right to claim. Legacy throughout is the upstream open-source engine GInfer was forked from, measured as it stood at the start of the campaign, before Quaestor rebuilt it.
Decode · long context
2.54×

C1 at 128K context

+154.0% over legacy
Prefill · long context
1.90×

C1 at 128K context

+89.7% over legacy
Decode under load
+62.3%
At C4
+53.5%
At C8

Decode aggregate, against the legacy engine

Double the load, still ahead
55.2
Tuned · C8
>
48.9
Legacy · C4

Decode per request, tok/s. Twice the concurrency, and each request still completes faster.

Long contextConcurrency 1 · 128K · RTX PRO 6000 Blackwell
MetricLegacy lowPost-tuning highImprovement
Prefill · PP614.214 tok/s1,165.255 tok/s1.90× · +89.7%
Decode106.743 tok/s271.109 tok/s2.54× · +154.0%
Internal measurement of the GInfer inference engine on an RTX PRO 6000 Blackwell, at 128K context. Columns compare the legacy engine’s low against the tuned engine’s high. Target model and quantization are to be stated with the full result set.
Concurrency scaling64K context · RTX PRO 6000 Blackwell · prefill and decode as requests go in flight
ConcurrencyPrefill · legacy → tuned · tok/sGainDecode aggregate · tok/sDecode per request · tok/sGain
C13,035.944 → 3,453.983+13.8%154.450 → 172.476same as aggregate+11.7%
C43,022.105 → 3,433.804+13.6%195.694 → 317.62848.924 → 79.407+62.3%
C83,017.146 → 3,394.819+12.5%287.486 → 441.37235.936 → 55.172+53.5%
Prefill gains are consistent and modest across concurrency, between 12.5% and 13.8%. Decode is where the tuning landed, and the gain grows once more than one request is in flight: 11.7% at C1, 62.3% at C4, 53.5% at C8. The practical consequence is on the fourth card above: at C8 the tuned engine delivers more decode per request than the legacy engine managed at C4, so concurrency can double and each request still completes faster. All figures on an RTX PRO 6000 Blackwell at 64K context, a shorter context than the long-context set above, which is why absolute prefill is higher here. Internal measurement, verification pending.
Across the fleet
Standard Benchmark, September 17, 2026

The tuned engine on every target it was built for: two models, six GPU configurations from SM80 to Blackwell, at one, four and eight requests in flight. The legacy comparison above shows how far the campaign moved one card; this is where the result stands on all of them.

Peak decode
3,224 t/s

Qwen 3.8 27B · RTX PRO 6000 ×2 · C8

Aggregate generation, TP2
Single card decode
2,483 t/s

Qwen 3.8 27B · RTX PRO 6000 · C8

One GPU, eight requests in flight
Cold prefill
34,514 t/s

Muse Glimmer 30B · RTX PRO 6000 · C1

Uncached prompt, computed in full
Time to first token
2.00 ms

Muse Glimmer 30B · RTX PRO 6000 ×2 · C1

Mean TTFT, cached wave
Qwen 3.8 27BC1C4C8CMP 170HX126346521RTX 3090149434614RTX 40903216871,128RTX 50906551,5972,144RTX PRO 60006581,7362,483RTX PRO 6000 ×29512,4433,224
Muse Glimmer 30BC1C4C8CMP 170HX132287422RTX 3090190371481RTX 40903607681,292RTX 50905591,3822,303RTX PRO 60005791,4422,431RTX PRO 6000 ×27671,8132,094
Decode throughput, aggregate tokens per second, at one, four and eight requests in flight. Every card scales with concurrency; the Blackwell parts scale furthest.
Qwen 3.8 27BSM80 · 86 · 89: AutoRound with Q4 DFlash2 · Blackwell: NVFP4 with NVFP4 DFlash2
GPUSMTPCCold PP t/sPP t/sTG t/sTTFT ms
CMP 170HX801C11,678774,281126.22.65
C41,6731,141,910345.77.17
C81,6721,244,242521.213.17
RTX 3090861C1857628,060149.03.26
C48631,071,410434.37.65
C88581,101,520614.514.87
RTX 4090891C12,007683,153321.03.00
C42,0041,132,717686.87.23
C82,003937,0591,127.917.48
RTX 50901201C18,204796,640654.62.57
C48,1971,052,6111,596.67.78
C88,1791,044,1752,143.715.69
RTX PRO 60001201C110,117915,598657.72.24
C410,0901,709,3511,735.74.79
C810,0511,560,7452,483.210.50
RTX PRO 6000 ×21202C18,253689,084951.12.97
C48,2531,448,0222,443.15.66
C88,2401,398,7603,224.311.71
Muse Glimmer 30BSame quantization per SM as above
GPUSMTPCCold PP t/sPP t/sTG t/sTTFT ms
CMP 170HX801C13,571716,054132.42.86
C43,5961,385,127287.25.91
C83,5641,048,915422.115.62
RTX 3090861C12,250688,968190.22.97
C42,2331,406,472371.45.82
C82,2121,285,274481.012.75
RTX 4090891C15,578690,922360.22.96
C45,5751,219,020767.86.72
C85,5711,401,7781,291.911.69
RTX 50901201C128,761680,102559.03.01
C428,543924,6631,382.18.86
C828,386850,4462,303.319.27
RTX PRO 60001201C134,514890,375579.42.30
C434,5761,583,7701,442.15.17
C834,4421,322,7522,431.312.39
RTX PRO 6000 ×21202C117,9831,024,483767.02.00
C417,9551,567,1321,813.25.23
C817,9801,514,3632,093.710.82
Standard Benchmark: 2,048 prompt tokens, 500 generated per request, one cold warmup wave and one cached measured wave per point. Cold PP is cold prefill computation. PP is aggregate logical prompt tokens divided by mean TTFT and includes prefix-cache reuse; it is not raw prefill compute speed. TG is aggregate decode across the requests in flight. Internal measurement, September 17, 2026.
The Standard Benchmark ships in GChat. Run it on your own hardware and publish the result to G.bench, the community results page.
Open G.bench →
01 · Output

The campaign produced a complete inference engine: runtime structure, the K/V system, per-SM hot paths aligned to each target's capability, per-SM quantization and calibration of the target models, the DFlash2 speculative path, and the tuned kernels beneath all of it.

It also produced a Quaestor Expert Model: a bespoke worker trained by the engine on its own output, preserved as weights and carried forward as an artifact of the engagement.

02 · Measurement

GInfer is measured against replay-verified baselines and frozen per-SM artifacts, under the workload profile agentic serving produces: high concurrency, long shared prefixes, repeated tool turns, and sustained high K/V occupancy. At 128K context and concurrency one, prefill reached 1.90× and decode 2.54× against the legacy engine. Under concurrency, decode aggregate gained 62.3% at C4 and 53.5% at C8.

These figures are internal measurement of GInfer, pending independent verification.

03 · Limits

One campaign, on one subject, does not establish that the method generalizes. It establishes that the machinery runs end to end on a hard, measurable problem and produces an auditable result, which is the narrower claim we are making.

The representation and architecture behind Tessera v1 remain undisclosed pending our initial patent filings. Nothing in this case study depends on them.

09What this demonstratesThe subject was a complete inference engine. The question was whether Quaestor could take a mandate that broad and return both a production engine and an expert in building one.
01

It designed the whole engine, and the custom kernels beneath it

The mandate was whole-engine: runtime architecture, per-SM hot paths matched to each SM’s hardware capability, the K/V system, and target quantization chosen for compute capability. Kernels came last.

02

The output is an artifact, not just an answer

Each pass preserves a set of trained weights. The expertise is retained rather than recomputed, and the next campaign in the domain starts from it.

03

Every gain is attributable

Frozen target models, isolated workers, one mechanism each, and cheap gates before expensive ones. A measured improvement belongs to a named change.

04

Nothing measured was wasted

Rejections carry as much training signal as wins, which is why the worker model improved while the kernels did.

The same engine, pointed at a different subject, runs the same campaign, and returns a different expert, with retained knowledge and expertise that make every future run faster and less expensive.
Quaestor · Sectile Research Laboratories

Quaestor solved our problem. It will solve yours.

Quaestor is available for custom research applications, scoped per engagement or via annual licenses — whether you require diligence, regulatory impact, contract and obligation review, portfolio assessment, M&A assessment, IT landscape assessment, complex migration planning, complex codebase reviews, compliance support, or any other domain where the analysis is broad, sourced, and will be contested. Quaestor will reduce your cost and produce your deliverables faster, with entirely defensible results.

Request a briefing Professional services
Design partners
partners@sectilelabs.ai
Press
press@sectilelabs.ai
Quaestor
About Quaestor →