Every way this company could be built, and what the evidence says about each.
Each dot is one of the 1,034 items the study produced; each line a relationship it found between them.
The thesis fails at verification, not at coverage
As proposed, the seed does not clear a fund's return bar. The coverage mechanism is real — peer-reviewed work puts the coverage-to-capability correlation at r = 0.95 — but business decisions have no executable verifier, and a formal ICML 2026 result shows a generation loop without one can only lose information. Trained capability has a 3–7 month half-life against 18–24 month procurement cycles, successor base models erase fine-tunes within a 2–4 month release interval, the finished training set is replicable for under $5,000, and the proof plan's judge panels collapse to roughly two effective votes. Of 347 pieces of evidence, 62% came back inconclusive, largely because the company has no public footprint.
What survives, across the 285 venture shapes generated and 46 taken to decision depth, is everything that imports a verifier or a payer the thesis lacks. Engineering programmes carry dated, externally recorded outcomes — schedule, cost, claims — so decision quality can be scored against reality rather than judged. A paid panel of thirty practising sales leaders replaces design-partner access that the evidence shows is legally blocked at PSU sites and priced at $100K–$2M elsewhere. Pilot contracts with a stated fee and trigger at signature replace access-for-evidence. Sealed third-party evaluation replaces self-grading. The machine is worth backing only inside one of those shapes, at a price well below the ~$15M ask.
Five numbers that set the price
Each is grounded against live sources; the findings and citations are below.
Eight findings that reprice the thesis
Each was produced by a web-grounded investigation, and each changes a specific number in the investment case. The headline is the conclusion; open a row for the evidence.
▶The generation engine's output cannot compound without a paid external verifier, moving the binding cost of the $6.5M generation plan from producing samples at $0.001–$0.01 each to scoring them at $0.50–$100+ each.1 / 8
A formal result at ICML 2026 (arXiv:2605.16379, Li, Sun & Deng) proves via the Data Processing Inequality that a generation-training loop with no external verification signal can only lose task-relevant information, making model collapse the predicted outcome. Mathematics and code have executable verifiers; business decisions do not, so the loop stays information-closed unless an outside signal is bought. The only demonstrated escape routes are expensive: expert judgment runs $0.50 to $100+ per sample against $0.001–$0.01 to generate one, and Bridgewater's specialist result required sustained expert-in-the-loop grading.
Source: arXiv HTML Version - An Information-Theoretic Criterion for Efficient Data Synthesis
▶If the bottleneck is reasoning depth rather than breadth of decision experience, the coverage machine addresses the wrong constraint and no volume of generated decisions repairs it.2 / 8
The deck's diagnosis, that business-decision capability is limited by breadth of decision experience in training, has the weakest empirical support of three competing explanations. Six independent studies show reasoning ceilings on complex strategic tasks that persist regardless of training-data breadth. The INFORMS Back Bay Battery simulation (Strategy Science, March 2026) found frontier models regressed on strategy performance even as their general benchmark scores rose. A second rival, absence of outcome feedback, drew moderate support from five further studies. The breadth hypothesis the company is built on came third.
Source: INFORMS Back Bay Battery Simulation (Strategy Science, March 2026)
▶The M2 proof, at its feasible budget, cannot statistically separate the three coverage-to-capability curve shapes it exists to distinguish, so a positive readout would not carry the evidential weight the tranche plan assigns it.3 / 8
At the 5,000–15,000 total expert judgments a $0.5M evaluation budget buys, the minimum detectable effect per coverage level is 5–10 percentage points. Distinguishing a shallow monotone curve of 3–5 points per increment from a plateau or from noise requires roughly four times that sample. Substituting automated judges does not close the gap: panels of language-model judges collapse to about two effective independent votes regardless of panel size (arXiv 2605.29800), and expert raters' self-agreement drops from ~0.85 to ~0.6 over four weeks without calibration anchors.
Source: Nine Judges, Two Effective Votes
▶The useful life of each trained specialist is one open-model release interval of 2–4 months, not the multi-year asset a $15M seed with deferred revenue implicitly prices.4 / 8
Open-model releases in the 20–40B parameter class arrive every 2–4 months, and successor base models routinely match or exceed prior specialist fine-tunes within a single interval. UpgradeBench (arXiv 2608.20918) measured specialist decay ranging from under one release cycle for text-to-SQL to about 14 months at best, and found adapter portability falls from 0.88–0.99 at short continued-pretraining distances to zero at 74% of a pretraining run. The Qwen3.6 to 3.8 step alone produced a 14-point jump that erased untargeted specialist gains. The only durable edge identified is data or workflow integration the base cannot acquire publicly.
Source: UpgradeBench: Studying Specialist Upgrades Across Consecutive Qwen Releases
▶The generated training data itself is worth under 0.1% of the $6.5M spent producing it to any follower, so the seed buys method and relationships, not a saleable data asset.5 / 8
At plausible training-set sizes of 10–100B tokens for a specialist decision model, a fast follower can regenerate an equivalent coverage-saturated synthetic set for $500–$5,000 within 18 months, using rented inference at current open-market rates. Even at 1T tokens the 18-month replication cost is roughly $35,000–$50,000, about 0.7% of the $6.5M generation budget. Separately, established model-extraction results show dense decision traces are easier to reconstruct from outputs than the thin confidence scores already shown sufficient for near-perfect model copying, so licensing the traces to labs would hand buyers the method.
Source: Together AI Inference Pricing
▶The engineering-programme validation domain loses both its data source and its exclusivity: the claimed NTPC sites cannot legally hand over records outside a public, non-exclusive process, and a replacement sourced at month 9 lands at month 18+, past the seed's useful life.6 / 8
CVC guidelines, Article 14 of the Constitution, and the Competition Act 2002 prohibit nomination-based awards of contracts or data access by Indian PSUs without competitive tendering, and post-award records are RTI-disclosable, so any competitor can obtain whatever a tender yields. The draft National Electricity Data Sharing Framework is voluntary and unsigned. No signatory or instrument could be confirmed for any of the four claimed NTPC sites. NTPC's standard procurement cycle runs 2–6 months with fixed enlistment windows, and no expedited data-sharing precedent by an Indian PSU was found.
Source: CVC Guidelines Booklet
▶Each design partner is a six-to-seven-figure cash cost rather than a free exchange, which raises burn across the M1–M5 sites and shortens the runway the $15M ask is meant to cover.7 / 8
The plan assumes partners supply decision records and expert hours in exchange for deferred capability. The market prices this the other way: Micro1, an AI data company at $100M ARR with a working product, pays enterprises $100K–$2M+ for workflow data access. Gartner's 2026 survey attributes 82% of AI startup valuation to proprietary dataset ownership, so enterprises increasingly treat their records as the asset. Standard design-partner exchanges include immediate tangible consideration — discounted access, roadmap influence, or cash — and free proofs of concept in regulated settings lack even a contract vehicle under which counsel can approve data release.
▶The coverage-to-capability link is now peer-reviewed fact in verifiable domains, which preserves the upside case: if the mechanism transfers to business judgment, the method is 150x more sample-efficient than volume-driven generation.8 / 8
The strongest support for the central claim: an ICML 2026 Oral (arXiv 2602.10388) shows Feature Activation Coverage correlates with downstream task performance at Pearson r = 0.95, and coverage-targeted synthesis reaches state-of-the-art results with 2,000 samples against MAGPIE's 300,000, across LLaMA, Mistral and Qwen. Adjacent results point the same way: PluRel (ICLR 2026) shows synthetic relational data obeying power-law scaling that transfers to real tasks, and REWIRE (COLM 2025) shows rewritten low-quality documents beating twice the filtered web data. All demonstrations, however, sit on benchmarks with measurable ground truth, not judged business decisions.
Source: Less is Enough: Synthesizing Diverse Data in LLM Feature Space with Sparse Autoencoders
Premises the study disproved
40 assumptions have been carried into the analysis and then invalidated against evidence. The 6 with the largest consequences are below. Confidence is the post-investigation figure — the lower it is, the harder the premise failed.
| Conf. | Premise that failed | What the evidence showed |
|---|---|---|
| 0.05 | Deciding the company's shape can be postponed until the M1–M3 evidence arrives; the optionality is free. | Failure to commit to a market is the best-documented startup killer: 43% of 431 VC-backed failures cited poor product-market fit, with a median 22 months from last fundraise to death. Hiring, partner contracts and compute commitments all differ by shape, so each deferred month builds the wrong thing. |
| 0.05 | A lab can buy the generated decision traces without being able to reconstruct the machine that made them. | Model-extraction research shows black-box access to thin outputs, even bare confidence scores, enables near-perfect reconstruction. The traces are far denser, encoding the discovered space, branching structure and stopping behaviour, and the named buyers are the one class with the research capacity to read them that way. |
| 0.08 | A partner's decision records are a stock that can be extracted once under a single consent. | Every regulatory framework examined, GDPR, DPDP and CCPA/CPRA, treats consent as decaying and requiring periodic refresh; enforcement actions use half-life language, and a $1.94B consent-renewal market has formed around the flow model. No credible source supports one-time persistent consent. |
| 0.10 | Model judges and expert human judges reach the same verdict, so cheap model-judged iteration is a faithful proxy. | In clinical reasoning, the closest tested analogue, a model judge approved 47.9% of incorrect outputs. Judges show self-preference, position and verbosity biases that flatter fluent fine-tuned output, and model-human agreement on subjective rubrics runs at r ≈ 0.27–0.32. |
| 0.15 | The rationale attached to a historical business decision reflects the actual reasoning closely enough that training on it teaches judgment. | Meeting minutes are sanitised accounts prepared for an audience; record fidelity decays within 24 hours of the decision. Three independent 2026 papers show training on post-hoc rationalisations teaches rationalisation, not causal reasoning, and the retrievable archive consists almost entirely of such records. |
| 0.20 | Business-decision capability is limited by breadth of decision experience in training, which the $6.5M generation and $7M training budgets exist to fix. | Of three competing diagnoses this had the weakest support. Six independent studies show reasoning ceilings on strategic tasks regardless of training breadth, including frontier models regressing on a strategy simulation, and five more support missing outcome feedback as the binding constraint instead. |
What survives, in six families
46 shapes have reached decision depth with quantified success criteria; 41 of them are grouped below into six theses about where value actually sits. The remaining 5 reached that depth after the grouping was made. The blue line under each shape is the study's own success threshold, not an estimate laid on top of it.
The evidence converges on one structural fact: the party being measured cannot own the measurement, and vendor-run evaluations are discounted by every audience that matters. No adopted business-judgment measure exists after five prior attempts, so the neutral scorer position is vacant. These shapes sell that neutrality itself, to funds, labs and buyers, and their value grows as machine-generated claims multiply.
Institutional access failed investigation: PSU records are legally blocked, employer consent is insufficient, and enterprise data agreements run 6–18 months. The decision structure the method needs lives in practitioners' heads, and individuals can be contracted in weeks at $40–$150 per hour. These shapes source the scarce input as a paid supply chain rather than a favour from a corporate sponsor.
Written decision records are post-hoc rationalisations with a 24-hour fidelity window, and organisations systematically discard the alternatives considered. These shapes create the record at decision time, or seal it before the outcome is known, so that hindsight cannot rewrite it. Capital programmes add the one thing the sales domain lacks: dated, externally recorded outcomes that a lender or insurer will pay to price against.
Free access-for-evidence trades fail structurally: a free pilot has no budget owner, no contract vehicle under which counsel can release data, and the market norm is enterprises being paid $100K–$2M for workflow data, not paying. These shapes replace goodwill with a signed instrument that names a fee, a trigger and a budget line at access time, so evidence and revenue come from the same document.
The cost advantage of a self-hosted mid-size model, roughly 18x per decision at sustained load, only matters on a workload that runs hundreds of times per decision. Rehearsal is that workload, and the need is already bought at high prices from humans in law, medicine and deal desks. The provable slice is narrow: predict the objections a named forum actually raises, without any training claim.
Enterprise security postures and shadow AI, present at 70% of companies, push serving inside the customer's boundary, where the vendor loses telemetry and the buyer loses verification. The operational layer that fixes this for both sides — sealed deployment, attested no-egress, in-perimeter evaluation with signed exportable scores — must be built anyway and is sellable to any vendor shipping into a regulated boundary, not only to Truna.
The cheapest ways to find out
Three tests that discriminate between the surviving families before committing real capital.
▶Build the adversarial frontier baseline$5–15K, two weeks
Give a disinterested contractor two weeks, the pilot records, and the newest frontier model to build an adversarial retrieval pipeline against the claimed ≤22% error framing. If the specialist's margin survives, the rehearsal and specialist shapes stay live; if it does not, every shape that sells trained capability dies and only the measurement and record-capture shapes remain. No published study settles this, so $15K buys an answer literature cannot supply.
▶Calibrate a judge panel first$20–40K, four to six weeks
Run the known-groups experiment: pay 8–10 practising sales leaders $100–200 per hour to blind-grade ~200 paired decisions containing planted quality differences and identical null pairs. If inter-rater agreement reaches κ ≥ 0.6, the panel, benchmark and sealed-evaluation shapes are viable and the proof plan can be powered properly; if experts cannot rank decisions consistently, no evaluation-led shape works and no training claim is measurable. The noise floor is currently unmeasured anywhere.
▶Sell three paid post-mortems~$10K founder time, six to eight weeks; fees offset cost
Offer a $5–25K retrospective decision post-mortem to arms-length prospects, excluding BarRaiser, with a signed statement of work and data-sharing clause. Three signatures settle whether enterprises will pay for decision work and release records under contract, reversing the access-for-evidence trade the evidence shows fails: free pilots produce no releasable records, and the market norm is enterprises being paid for data. Each engagement also measures whether released records contain any rejected alternatives.
What was actually searched
Every conclusion in the verdict traces to a logged investigation. These are the totals behind them.
Evidence verdicts across 347 grounded investigations
62% came back inconclusive and 40 actively contradicted the claim under test. A study that mostly confirmed itself would be the warning sign.
Value axis: count of evidence items (347 total). Category axis: verdict recorded by the grounding worker. Source: run artifact 07_evidence.json.
The axes of the search
33 distinct analytical frames were applied to the venture. This is the enumerated set — and the honest read is that 6 were pursued hard (highlighted) while the rest were opened once. The gap is the study's remaining work, not a claim of completeness.
The corpus
| Artifact | Count | State |
|---|---|---|
| Canvas nodes | 1,034 | 1,610 edges, depth 0–3, 190 distinct node types |
| Venture shapes enumerated | 285 | 46 taken to decision depth with KPI thresholds and named kill conditions |
| Open questions | 472 | 217 resolved against live sources, 245 still open, 6 in flight |
| Assumptions | 271 | 61 grounded, 40 invalidated outright, 19 contested where the research did not settle them, 142 not yet tested |
| Evidence items | 347 | 316 carry a retrievable URL; 143 at confidence ≥0.8 |
| Risks / controls | 290 / 440 | Every control names its target risks, an effectiveness estimate and a decay rate |
Negative space
What was looked at and rejected, which matters more than what was kept — it proves a threshold existed.
| Score | Signal |
|---|---|
| 0.45 | ICML position paper supports need for measurable information gain in synthetic data |
| 0.45 | Harness-of-Harness: Multi-Day Autonomous Software Development |
| 0.40 | ICLR 2026 paper on evaluating LLM forecasters relevant to Truna's evaluation design |
| 0.40 | ExToken bridges coverage and VLA fine-tuning with behavioral priors |
| 0.40 | No domain extension or adoption of the PA policy-retrieval approach in 5 months post-publi |
| 0.40 | ProbMoE: TPMs as differentiable surrogates for discrete expert routing (ICML 2026) |
| Source | Why it stopped |
|---|---|
| arXiv — cs.AI / cs.CL / cs.DB / cs.LG | Went quiet after 2 probes returned nothing new; 6 signals accepted from 3 scans before that |
| VynFi / DataSynth (github.com/mivertowski/syntheticdata) | Went quiet after 2 probes returned nothing new; 3 signals accepted from 3 scans before that |
| Twinning Labs / Posterior Twins (arXiv:2606.16415) | Went quiet after 2 probes returned nothing new; 3 signals accepted from 3 scans before that |
| [core] Small-Data Large-Scale Decision Optimization group (… | Went quiet after 2 probes returned nothing new; 5 signals accepted from 3 scans before that |
| [adjacent] Coverage-Guided Exploration in Model-Based RL (I… | Went quiet after 2 probes returned nothing new; 3 signals accepted from 3 scans before that |
| [adjacent] TalkingTrees / Interpretable Decision-from-LLM g… | Went quiet after 2 probes returned nothing new; 1 signals accepted from 3 scans before that |
30 sources are under active watch — 16 in the core field and 14 in adjacent fields, each of the latter required to name what transfers. 12 have gone quiet, meaning recent probes surfaced nothing new; they stay in rotation at lower priority rather than being dropped. No source has been retired for producing noise.
What this study does not establish
Volunteered rather than discovered, because the gaps change how much weight the rest carries.
- Truna itself was largely unreachable: the company operates in stealth with a single landing page, so its coverage-versus-compute curve, evaluation criteria, per-arm training cost, partner agreements, cap table and founder identities all remained unverified, and every company-specific claim in the deck still rests on the deck.
- Of 472 questions raised, 217 were resolved, and 62% of the evidence came back inconclusive, much of it because sustained HTTP 429 rate-limiting blocked web search across whole sessions and left only directly fetched pages.
- The two measurements the decision most needs do not exist anywhere: no published figure for expert inter-rater agreement on paired business decisions, and no published comparison of a specialist against an adversarial frontier-plus-retrieval baseline; both require commissioned experiments, not reading.
- The claimed design-partner relationships with NTPC and a pharma manufacturer could not be confirmed as instruments rather than conversations, and the named-approver, budget-calendar and record-format questions at every site were answerable only by asking the partners directly.
- Deal-side facts stayed dark: which funds have seen deck v8, whether the round is contested, the lead fund's reserves and tranche conduct, and the syndicate's verification capacity are all private records no search could reach.
- Of 285 venture shapes generated, 46 were taken to decision depth; the remainder carry design-stage criteria only, and none has been tested against a paying counterparty.
Open questions, ranked by whether the answer changes the decision
245 remain open. These are the eight the study itself ranks highest, and they are the natural first work order for diligence.
| Priority | Question |
|---|---|
| 1.00 | Whether any measured M2 lift is new capability or recovered generator capability — nothing in the disclosed baseline set (frontier, RAG-agent, size-matched open) answers this |
| 0.96 | How much better the coverage-trained specialist is than the same untuned open base model, run on identical hardware at identical cost — the only comparison that isolates what Truna adds |
| 0.96 | How different the outcome distribution of voluntarily contributed programmes is from the unselected public population |
| 0.95 | Which specific component of the pipeline the founders believe survives publication — the one thing a well-resourced lab reading the paper still could not do |
| 0.95 | How much does the frontier-coverage gap move when the concept granularity is redefined in three defensible alternative ways? |
| 0.95 | What a frontier-equivalent 200-way deliberation on one business decision costs today, and what the four-quarter trend and bundling trajectory of reasoning-budget tiers looks like |
| 0.94 | What share of any back-tested discrimination comes from documentation behaviour rather than from decision content |
| 0.93 | For each target sponsor role, what data access can that role authorise alone, and where does the approval chain for CRM export, mail connectors and DPAs actually terminate? |
Priority is the study's own ranking on a 0–1 scale. Source: run artifact 04_unknowns.json.
This view reorganises the whole study around the eight questions the investment turns on, so each claim sits next to the evidence for and against it and the questions still open. Items are matched to questions by keyword, so the counts are reproducible rather than hand-picked. Nothing is left out — the other tab has all 1,034 items.
Q1Is missing decision experience actually the constraint the $13.5M plan attacks?
Contested. Of three rival diagnoses this has the weakest support: 6 studies show reasoning ceilings regardless of data breadth, 5 support absent outcome feedback, and the coverage-to-capability link is proven only in verifiable domains.
Q2Can business-decision quality be measured without an executable verifier?
Unproven. An ICML 2026 proof shows synthetic loops degrade without external signals; judge panels collapse to ~2 effective votes; expert inter-rater agreement on paired decisions is unmeasured; every documented synthetic-data success had automatic ground truth.
Q3Can the claimed enterprise decision records actually be obtained?
No, not as planned. A PSU cannot legally grant nomination-based data access; employer consent is insufficient under DPDP and GDPR; the market pays enterprises $100K–$2M for workflow data; free pilots produce no releasable records.
Q4Do the records that survive contain judgment worth training on?
Doubtful. Written rationales are post-hoc accounts with a 24-hour fidelity window; three 2026 papers show training on them teaches rationalisation; minutes structurally exclude alternatives; compliance review strips the contested content.
Q5Does a trained specialist stay ahead of frontier and successor models?
Briefly. Open base releases arrive every 2–4 months and untuned successors erase fine-tune advantages within one interval; trained capability decays with 3–7 month half-lives against 18–24 month procurement cycles; test-time compute substitutes for specialist weights.
Q6Is the $6.5M generation budget buying a durable asset?
No. A follower can regenerate the coverage-saturated training set for under $5,000; model-extraction results show dense traces hand the method to any lab buyer; and take-or-pay GPU reservations survive a month-9 kill decision.
Q7Will anyone believe an evaluation the equity holders design and run?
No, as designed. Vendor-run evaluations are discounted for file-drawer and funding bias; the fastest access site is founder-owned and related-party; a three-founder team cannot hold the training-evaluation firewall; external custody is the standard remedy.
Q8Is anything here worth backing, and on what terms?
Conditionally. The coverage mechanism has peer-reviewed support (r = 0.95 in verifiable domains); the case survives only in shapes with an external verifier, paid contracted pilots and sealed evaluation, priced well below the ~$15M ask with pre-registered milestones.
i
How this is organised. Everything sits in one of three sections, listed in the column on the leftunder Browse. Pick one and its contents fill this panel; click any row to go down a level. The trail at the top always shows where you are and takes you back. 985 of the 1,034 items have exactly one parent and nothing goes more than three levels deep, so it reads as an outline.
The maps. Any item can be switched to Map to see just its own neighbourhood — what it sits under, what it breaks into, and what links to it from elsewhere. The Map button in the left columnunder Browse instead shows the whole thing in three dimensions: either the 625 cross-links on their own, which is the part this outline cannot show, or all 1,034 items at once.
The other 625 connections do not follow the parent-child structure — they are things like one item feeding a risk into another (risk_propagation 228 · dependency 187 · causal 142 · financial_flow 53 · control_mitigation 15). They show as dashed gold lines on the map and as “feeds into / fed from” on the page.
Why the counts differ. The structure holds 1,034 items. The 285 venture shapes, 271 assumptions, 472 questions, 290 risks and 440 controls are not items — they sit alongside or hang off them, which is why a tag never finds them. Use the buttons above the tags to browse each set on its own.
Loose findings is a remainder, not a category: roots that never attached to the venture or to the outside world. Mostly scout signals that did not land under their parent, plus a few strays. It is named for what it is rather than dressed up as something coherent.
Tags versus groups. Tags are the exact kind the machine assigned (190 of them, so the long tail is behind “show all”). Groups bundle those into 13 broader buckets.
What is thin. 452 items have a written explanation; the other 582 are a headline only, because the run was stopped before every branch was written up.