YATHARTH Defence vision integrity// 4 runs
Prototype · demoPrototype demo - results may contain errors. These checks were run ahead of time on a few datasets. A live assessment takes around 15 minutes per run of about 1,000 images and needs a dedicated GPU (these ran on a 6 GB RTX 4050 laptop GPU), so this page is the run history: a replay of a previous run, showing exactly how it was processed. See the four runs ↓

Every file,
taken apart.

From pixels to provenance.
Multi-axis analysis for trustworthy
computer vision.

Scroll to take it apart ↓
YATHARTHIntegrity across contributorsDataModelsInferenceSame image. Seven perspectives. A clearer truth.
Run it

What it takes to run.

Everything runs on one machine, offline. These are the figures from the machine that produced the runs on this page.

01

Hardware

  • An NVIDIA GPU with 6 GB of memory or more (tested: RTX 4050 laptop GPU)
  • 16 GB of RAM
  • About 4.4 GB of model weights (DINOv2 and Qwen3-VL-4B), stored locally
02

Time

  • Around 15 minutes per run of about 1,000 images
  • Longer the first time the vision model sees a dataset; its answers are cached after that
  • The replays on this page play back a finished run in about a minute
03

Built to be checked

  • Every flag carries the measurement behind it, in declared units
  • Confidence is measured against known answers, not assumed
  • No data leaves the machine, and each run is fingerprinted in a tamper-evident log
Does "quarantined" mean the file is malicious?

No. It means independent checks agree that something is off, so a person should look before the file is used. The file is kept, unchanged.

Can it tell bad weather from tampering?

That is what the drift check is built for: weather changes every source and class and is explained by brightness, blur and colour; tampering is usually narrow and unexplained. So far it has been tested on simulated conditions.

Is this real time?

No. Each assessment is a batch run of about 15 minutes. This site replays runs recorded earlier.

Does any data leave the machine?

No. Every model runs locally and offline, and every page here is self-contained.

Verify

4 datasets. One pipeline.
Every decision shown.

Each run went through the same checks, under the same rules. See the pipeline run to watch every check work through every image; read the report for the evidence behind the decision; download the exact data and models it assessed.

01

Synthetic showcase

A generated dataset with every attack planted on purpose: label flips, duplicates, flooding, train/test leakage, foreign images, triggers, a poisoned model and a tampered prediction log.

DecisionQUARANTINEDo not use as submitted: something it must guarantee was clearly broken.
302images
302objects
882findings
44kinds
0readiness
Most frequent findings
  • witness scene mismatch302
  • witness label disagreement270
  • swapped label58
  • model spectral signature45
  • spectral signature44

Access L3 · reference R1 · backbone stub-pixelstats · policy 2026.09.28.3

02

Real dataset, attacked

A real military-vehicle dataset with attacks planted into it, submitted together with a model trained on it that carries a backdoor.

DecisionQUARANTINEDo not use as submitted: something it must guarantee was clearly broken.
1229images
1417objects
1604findings
43kinds
0readiness
Most frequent findings
  • witness scene mismatch444
  • witness marking324
  • witness label disagreement297
  • swapped label212
  • model spectral signature148

Access L3 · reference R1 · backbone dinov2-small · policy 2026.09.28.3

03

Real dataset, as supplied

The same real dataset, untouched, with no model: what the assessment says about data nobody tampered with.

DecisionQUARANTINEDo not use as submitted: something it must guarantee was clearly broken.
1000images
1120objects
780findings
8kinds
0readiness
Most frequent findings
  • witness scene mismatch365
  • witness marking169
  • swapped label145
  • witness label disagreement87
  • trigger texture8

Access L3 · reference R1 · backbone dinov2-small · policy 2026.09.28.3

04

Real dataset, after curation

The previous run after removing what it quarantined and assessing what remained, until nothing more was quarantined.

DecisionREVIEWNothing clearly broken, but a person should check the flagged items first.
890images
988objects
545findings
9kinds
27readiness
Most frequent findings
  • witness scene mismatch301
  • witness marking133
  • swapped label64
  • witness label disagreement33
  • trigger texture5

Access L3 · reference R1 · backbone dinov2-small · policy 2026.09.28.3

Explore

More context, coming next.

Deeper pages for anyone who wants to check the reasoning, not just the result. They are being written now and will be linked here.

In preparation
A

How each check works

The twenty checks one by one: what each measures, in what units, and what it cannot see.

In preparation
B

Dataset history

How one dataset changes from run to run: what was fixed, what is new, and what drifted as fresh data came in.

In preparation
C

Stress tests

Where it works and where it does not: weather, new sensors, look-alike classes and attacks, each with its measured result.

Trace

The inputs, exactly as assessed.

Images, annotations, models, sealed prediction logs and the ground truth each run is scored against. Every archive carries a README and a SHA-256 for every file inside it.

RunImagesModel filesFilesSizeSHA-256 of the zip
Synthetic showcase61266343.5 MB7dc0c8b8c958b1d1…Download ↓
Real dataset, attacked189991925100.5 MBd404755e5984d918…Download ↓
Real dataset, as supplied16000161171.1 MBe28b6fe16086fdd4…Download ↓
Real dataset, after curation16000161171.1 MBc60831cf6656cab5…Download ↓

The synthetic archive contains deliberately hostile model files (a pickle that references os.system, an ONNX file with a Python operator, one that reaches outside itself). They exist so the safety gate has something to catch; Yatharth parses them without loading them. Do not load them yourself.

Assure

Facts, evidence, judgement - kept apart.

A single risk score hides which check fired and why. Yatharth keeps the three steps separate, so every decision can be traced back to the measurements under it.

01

Facts

Ingest records what each file is: hashes, embeddings, box geometry, who contributed it. It never judges.

02

Evidence

Twenty engines each report a score in declared units, with the material behind it and its limitations. No engine holds a threshold. A check that cannot run says so, and is never counted as a pass.

03

Judgement

Only the correlation layer, reading the policy file, sets severity, confidence and what to do. Quarantine needs independent kinds of evidence to agree; confidence is measured against planted ground truth.