Assessment method · v3.0

Verification Methodology

How FactNot turns available evidence into an inspectable assessment. A score is not a guarantee of truth, and every historical run keeps its own version and weights.

01

The current formula

Version v3.0 combines four bounded components into a 0–100 assessment after a checkworthiness gate. Checkworthiness describes whether a claim is testable; it is never a truth score.

  • With evidence agreement: LVF 60%, NLI 20%, source reliability 10%, checkworthiness 10%.
  • When NLI is not run: LVF 75%, source reliability 15%, checkworthiness 10%.
  • Claims below the 30% checkworthiness gate are skipped before the costly assessment stages.
02

Six LVF criteria

Independent assessors score six criteria from 0 to 10. The reconciled, weighted result becomes the language-verification-framework component.

  • Source provenance (20%): Directness, traceability, and quality of the cited source record.
  • Corroboration (25%): Agreement from independent evidence rather than repeated copies of one claim.
  • Internal consistency (15%): Whether the claim and evidence remain coherent without material contradictions.
  • Specificity (10%): Whether the claim is precise enough to test against evidence.
  • Temporal validity (15%): Whether the evidence is current enough for a time-sensitive claim.
  • Bias signal (15%): Whether framing or source incentives weaken confidence in the assessment.
03

What labels mean

The status describes the verification outcome; the High, Medium, and Low display bands are a compact reading aid. They do not replace the number, checked date, methodology version, or limitations.

  • Factual: 60–100. Rumour: 40–59. Opinion: 0–39.
  • Display bands — High: 91–100. Medium: 60–90. Low: 0–59.
  • Supported language is an assessment label, not a warranty or professional conclusion.
04

Provenance and independent support

Good corroboration depends on independent origins. Ten articles repeating one wire report do not count as ten independent confirmations.

  • Source provenance asks whether a source is direct, traceable, and appropriate for the claim.
  • Publisher reliability is a coarse supporting signal and cannot rescue weak claim-level evidence.
  • Public score summaries expose counts and component results, never private source bodies, excerpts, prompts, model responses, or reviewer notes.
05

Automation and human review

Automated stages classify claims, gather evidence, compare passages, score criteria, reconcile assessors, and run a fail-closed quality review before a fact is updated.

  • AI assessors can be wrong, inconsistent, or overconfident.
  • Risky, sparse, or contradictory results can be held for more evidence or human review.
  • Public votes and reactions affect review priority only; they do not alter new truth scores.
06

Freshness and unavailable checks

A verification describes the evidence available at its checked date. Time-sensitive facts can become stale and should be rechecked.

  • When external search is disabled or unavailable, the result says not run rather than treating missing search as evidence against the claim.
  • When NLI is unavailable, the fallback formula is used and the component says not run.
  • A later run creates a new explanation; older runs retain their persisted version, weights, and component values.
07

Challenges and evidence feedback

A factual challenge is a structured request for review, separate from an abuse report. Submitting evidence does not immediately publish it or change a score.

  • The submission is validated, rate-limited, and queued for review.
  • A reviewer can request more information, reject the challenge, or accept evidence for a later verification run.
  • Accepted evidence can inform a new versioned assessment; the original historical run remains unchanged.
08

Known limitations

Verification quality depends on claim specificity, evidence access, source independence, freshness, and model behavior. Important decisions still require primary-source review.

  • A score is an assessment from available evidence, not a guarantee of truth.
  • Repeated reporting may trace to one origin and must not be counted as independent corroboration.
  • This run found fewer than two source records, so independent corroboration is limited.
  • External search can be unavailable or disabled; that state is reported as not run, not negative evidence.
  • Evidence agreement can be unavailable; fallback weights are used and NLI is reported as not run.
  • Time-sensitive claims can become stale and need another verification run.
  • Automated assessors can be wrong; risky or disputed results may require human review.
09

Version changelog

Each completed run stores its methodology version and exact weights so the displayed explanation remains faithful after the current method changes.

  • v3.0: Public votes and reactions became review-priority signals and no longer alter truth scores.
  • v3.0: Public score and history surfaces use visibility-checked, allowlisted projections.
  • v2.0: Introduced persisted component breakdowns, run weights, NLI support, and methodology versioning.
  • v2.0: Historical v2.0 records continue to render with the weights stored on each run.