M
O
Open Dashboard arrow-right icon arrow-right icon
Layered blush pink mountain ridge
Faceted pink mountain ridge
Saturated pink foreground terrain

Version 1.0 · October 2026

Whitepaper

The design and limits of an
evidence-bound AI jury.

Explore this document ↓

01 / WHITEPAPER

Abstract

MOCHI checks a factual claim against user-submitted evidence with a jury of AI models. The launch service combines browser encryption, remotely attested confidential workloads, independent model-family votes and an on-chain settlement record. A verdict requires three valid matching answers; otherwise the round is HUNG.

The private answer is returned as ciphertext and decrypted on the requester’s device. Its hash is checked against the chain record. Evidence spans, signatures and receipts support inspection of execution. These mechanisms bind the record to the computation; they do not prove that a claim is true, a source authentic or a model’s reasoning sound.

02 / WHITEPAPER

The problem

A claim can circulate faster than its evidence. A short quotation can omit a date, qualification or conflicting passage. A link alone does not reveal which text was used to reach an answer, and a confident model response can blur the distinction between supplied evidence and prior knowledge.

A single model also gives the requester little visibility into alternative readings. Repeating the same model is not necessarily independent scrutiny. MOCHI makes the submitted evidence explicit, asks different model families the same bounded question and exposes failure to reach consensus as a protocol outcome.

The launch service evaluates the evidence package the requester provides. It does not search the web, fetch source URLs or certify the identity of the source publisher.

03 / WHITEPAPER

Design goals

  • Evidence-bound answers. Every counted answer must cite an exact passage in the submitted text.
  • Explicit disagreement. A failed consensus remains HUNG; a majority is not presented as settled at launch.
  • Private content. Encrypt before upload, confine plaintext processing to approved workloads and encrypt the result to the requester.
  • Inspectable execution. Bind workload identities, votes, answer hashes and signed receipts to protocol records.
  • Defined payment outcomes. Quote before payment and settle responding seats, protocol fees, timeouts and expiry according to the escrow contract.
  • Honest boundaries. Separate measured evaluation results, live deployment configuration and future capabilities.

04 / WHITEPAPER

Architecture

The browser verifies the intake workload and encrypts a deterministic claim-and-evidence package to its attested key. The intake enclave validates the package, computes its commitment and input tier, stores sealed content and signs intake provenance. The requester then opens a private USDG-funded query on chain.

The orchestrator observes the query, selects the protocol’s registered seats and coordinates encrypted dispatch. Intake seals the document separately to the jurors and sends a seed to the consensus enclave. Each juror enclave service obtains confidential inference, validates the structured response and exact evidence quotes, and signs a vote.

From browser encryption to a private, verifiable resultThe browser encrypts the claim. Attested intake sends encrypted packages through the orchestrator to three independent juror enclaves using confidential inference. Attested consensus requires all three to agree. The chain records verdict commitments and receipts; the browser decrypts the private answer and matches its hash.Browser encryptionClaim + source excerpts · local recovery keyAttested intake enclavePinned identity · sealed storage · encrypted dispatchOrchestratorSchedules work · routes ciphertext · submits transactionsQwen 3.6Gemma 4GPT-OSS 120BJuror serviceJuror serviceJuror serviceAttested confidential inference · signed votesAttested consensus enclave3 / 3 valid agreement → verdict · otherwise HUNGOn-chain verdict + answer hashSigned receipts · Merkle-root anchorsPrivate ciphertext → browser decryption + hash check
Content moves as encrypted envelopes outside approved workloads. Launch intake, juror and consensus services share one confidential VM with separate role keys. The three model families are independent selections, not independent hardware or operators. The orchestrator coordinates execution without plaintext claim content.

The consensus enclave collects the valid responses, applies the agreement rule and signs the round decision. The orchestrator submits that decision and juror votes to MochiVerdicts. The chain records the verdict status, answer hash and evidence and attestation commitments; QueryEscrow settles the corresponding payment.

The private result is sealed to the requester’s result public key. The browser retrieves the ciphertext, decrypts it locally and matches the answer hash to the chain. Signed receipts can be verified and checked for inclusion in a published ReceiptAnchor root. A signature check and an on-chain anchor check are distinct operations.

05 / WHITEPAPER

Jury design and model selection

N3 uses three seat classes: LARGE_A, DOC_SPECIALIST and DISSENTER. The launch mapping is Qwen 3.6 (qwen/qwen3.6-35b-a3b), Gemma 4 (google/gemma-4-31b-it) and GPT-OSS 120B (openai/gpt-oss-120b), respectively.

Class names describe assigned roles, not a guarantee of expertise or opposition. Three model families reduce dependence on a duplicated model vote, but they may still share training data, biases and failure modes. Infrastructure and inference-provider dependencies remain shared.

The 60-claim SciFact evaluation

The internal evaluation sampled 60 SciFact development claims: 20 support, 20 contradict and 20 not-enough-information cases. It used cited abstracts packaged through the SDK, the production FREEFORM_FACT prompt, strict output validation, attested Phala ACI inference and the production consensus rule.

Launch panel results on the internal sample
Correct39 of 60
Wrong7 of 60
Unresolved14 of 60
Median time22 seconds
Mean inference cost$0.016 per evaluated panel

Two duplicate Gemma seats agreed on 59 of 60 claims, motivating a third family rather than another copy of that model. Two evidence-handling fixes also mattered: requiring a quote for insufficient-evidence answers, and allowing one repair request for an inexact quote while retaining verbatim validation.

Seven wrong outcomes still passed unanimity. The sample is small, scientific and not representative of general public or crypto claims. Inference cost is not total service cost, and evaluation latency is not a launch SLA. Read the model-selection report for methodology and comparisons.

06 / WHITEPAPER

Consensus and evidence rules

The claim-review adapter asks for exactly one of supported, contradicted, missing_context or insufficient_evidence, based only on the submitted evidence. It treats the claim, metadata and excerpts as untrusted data rather than instructions.

The protocol’s threshold is ceil(3N / 4); at launch N = 3, so all three seats must agree. A value without valid evidence spans does not count. A timeout or malformed response is not a vote. Any failure to reach the required agreement yields HUNG and no claim outcome.

Jurors must quote at least one exact passage for every answer. For insufficient evidence, they quote the passage closest to the claim to show what the material does and does not establish. Span validation binds the quote to the uploaded text; it does not authenticate the source or validate the logical inference.

Unanimity is a conservative resolution rule with an availability cost. It suppresses majority answers when another juror disagrees or fails, but it cannot eliminate shared mistakes or prompt-injection risk.

07 / WHITEPAPER

Confidential computing and attestation

The launch runtime uses Intel TDX confidential VMs managed through Phala dstack. Remote attestation binds a workload measurement to its signing and encryption identities. The browser rejects an intake identity or measurement that differs from the pinned public configuration and checks quote freshness and verification collateral before sending encrypted content.

JurorRegistry tracks approved measurements, keys, classes and active attestation status. The intake, juror and consensus services exchange encrypted envelopes addressed to registered workload identities. The confidential inference path verifies Phala ACI workload reports and receipts binding the model and request/response exchange.

Attestation establishes evidence of a measured execution environment. It does not prove that the measured code is bug-free or that a model is correct. The trust boundary includes Intel’s hardware and verification chain, dstack and its key-management behavior, approved workload measurements, confidential inference verification and the client software serving the page.

Operators can deploy workloads and affect scheduling and availability. Governance can approve new measurements and identities. Those powers remain relevant even when operators cannot read protected workload memory.

08 / WHITEPAPER

On-chain protocol

A paid query progresses from OPEN to SEALED, then to DECIDED on a verdict or HUNG on failed consensus. QueryEscrow collects USDG and calculates class-based fees. JurorRegistry supplies eligible identities; MochiVerdicts checks signatures and stores the decision commitments; ReceiptAnchor records receipt roots.

The query deadline is one hour after opening under the launch TTL. An unanswered OPEN or SEALED query can be moved to EXPIRED by an on-chain expiry transaction after the deadline, refunding the remaining escrow. Time alone does not trigger a transfer. The orchestrator can perform expiry while running; a stalled service does not remove the need for a transaction.

HUNG settles the responding seats and returns the protocol fee plus timed-out seat fees. It is distinct from an unanswered query and cannot use that expiry path. An insufficient-evidence verdict is a DECIDED, chargeable outcome.

Governance is behind a timelock with a 60-second launch delay. The controller can change its delay; deployment tooling bounds are not an immutable contract limit. A guardian can pause new openings, while settlement and payouts remain available; governance controls unpausing.

Mainnet addresses are published at launch. The live website configuration is /mochi-config.json. Consult the published deployment record for the full contract set. The initial jury is team-operated, and the protocol does not claim an open decentralized operator network.

09 / WHITEPAPER

Economics

The configured short N3 tariff is $0.10 USDG at the minimum input tier, before requester network gas. It allocates $0.08 to juror/operator fees and $0.02 to the protocol. These are tariff allocations, not measurements of each model’s marginal cost.

Short N3 quote at tokensK = 1
Seat or allocationUSDG
Qwen · LARGE_A$0.028
Gemma · DOC_SPECIALIST$0.030
GPT-OSS · DISSENTER$0.022
Protocol fee$0.020
Total before gas$0.100

For each class, the contract adds its base fee and per-tier fee multiplied by tokensK. Intake estimates the tier from extracted text length. The protocol fee is the greater of the configured minimum and 20% of juror fees. Larger inputs raise the quote; the live contract quote governs payment.

On a successful verdict, governance can reserve a share of the protocol fee for future human panels. At the contract default of 25%, that is $0.005 of the short tariff, leaving $0.015 for staking or the configured review recipient. Human escalation is off at launch, and the launch configuration may set the reserve share to zero. On HUNG, responding jurors are paid and the protocol fee and timed-out seats’ fees are refunded. Gas is separate.

MOCHI is issued by the team. Its contract address will be published; this paper does not announce a token price, supply, sale or return. Customers pay in USDG and do not need MOCHI. A separate tokenomics proposal, which is not active, describes using eligible service surplus for purchases and manual burns after operating obligations and reserves. Activation is separate; a check does not automatically purchase or burn tokens.

Pricing comes from the launch tariff configuration and settlement from QueryEscrow. Hosting, failed inference, settlement gas and other operating costs are not fully captured by the evaluation’s inference-cost figure.

10 / WHITEPAPER

Privacy

For the private paid path, the claim and submitted excerpts are encrypted in the browser before upload. Plaintext exists inside the approved processing workloads; persistent private content is sealed ciphertext. On-chain and database records retain hashes, commitments, signatures and operational metadata rather than plaintext private claims or answers.

Public chain data still reveals wallet activity, payments, query identifiers and timing. Operators can observe traffic sizes and service availability. Production telemetry is restricted to content-free timing and status fields, including query and model identifiers; it is not a log of private prompts or results.

The requester keeps the result key locally and in the private recovery file. Losing the key can make the private result inaccessible; disclosing it can expose the result. The separate invitation research flow processes public claims through external providers and is labeled unattested research. Its share and retention behavior should not be treated as the private protocol’s guarantees.

11 / WHITEPAPER

Limitations and risks

  • Models can be wrong together. Unanimity and exact quotations do not establish truth, source authenticity or complete context.
  • Evidence can be incomplete or adversarial. Submitted URLs are not fetched or independently authenticated. Prompt and span checks do not eliminate all malicious-input risks.
  • HUNG costs and availability. Failed consensus produces no answer, while responding jurors still earn their fees. Timeouts, retries and outages can delay results.
  • Provider dependencies. The launch model families share confidential inference infrastructure and depend on provider availability and verification services.
  • TEE assumptions. Hardware, firmware, attestation collateral, key management and measured software may contain vulnerabilities. Governance controls approved measurements.
  • Contract and governance risk. Bugs, configuration changes, a short timelock and privileged pause powers affect the service. Network gas and transaction failures remain requester concerns.
  • Client and key risk. A compromised browser can expose content before encryption. A lost or disclosed recovery file affects result access and privacy.
  • Scope. The launch service provides evidence-based claim assessments. It is not financial or legal advice and is not a substitute for professional review.

12 / WHITEPAPER

Roadmap

Launch is limited to the N3 private claim check. Human panel escalation, appeals and larger juries are off. Broader mechanisms in the contracts or SDK remain outside the launch offering.

The roadmap orders work around core jury execution, verifiable records, standing feeds and builder access. Future feed expansion, agent access and human-panel escalation depend on validation; human panels are not available at launch. The roadmap commits no release dates or production capacity.

Attestation rejection, privacy and ownership controls, timeout recovery, HUNG settlement, expiry refunds and stale-feed behavior remain validation gates as the service expands. Token purchase and burn activation follows separate activation conditions, not a promised roadmap date.