Cultural Atlas
Indian tourism references

Culture: how it works.
Follow one scene through the benchmark.
Wikipedia · UNESCO
Sahapedia · IGNCA
Wikidata · Wikimedia Commons · ASI
Census of India · Ministry of Tribal Affairs
Magh Bihu dawn near Nagaon, Assam: a bamboo-and-hay meji bonfire, a bhelaghar hut beside it, and women offering til pitha and coconut laru.
Start with the named sources above. Link each cultural claim to its specific supporting reference.
Claude Code helps turn the scene into a prompt and required visual checks.
Validation gates + Gemma 4 31B check the claims against their evidence.
Experts review cultural accuracy and translations before test-set release.
Generate images from approved prompts and check each required detail. Planned.
Are the fire, hut and offerings correctly identified as Magh Bihu? Does the translation preserve those details? Machine verification alone does not release the scene.

Culture: what we have.
The data behind the example.
authored scenes
target scenes
categories
states / UTs
*Machine-verified. All 237 current scene records remain drafts.
Pending: expert checks, 1,356 translation rows and 263 more scenes to reach 500. Five-language design: EN / HI / TA / TE / BN.
Magh Bihu · Assam

Regional clothing reference; this photo does not depict the Magh Bihu ritual or a scored output.

Culture: two kinds of expectation.
Explicit details. Implicit meaning.
“Magh Bihu in Assam: villagers gather at dawn around a bamboo-and-hay meji bonfire.”
Did it show what we asked?
Details written directly in the prompt.
Example checks: villagers, dawn and the bamboo-and-hay meji bonfire.
A vision-language model inspects the image against a checklist. People verify mistakes and uncertain judgments.
Does the culture make sense?
Expectations inferred from the cultural context, even when unstated.
Example checks: does the ritual fit Magh Bihu? Are clothing and offerings appropriate to this Assamese scene?
Review meaning, regional context and stereotypes against evidence. Record valid local variations.
0 no alignment · 0.5 partial · 1 complete
One alignment rating; below 1, flag explicit / implicit / both and explain why.
Our proposed workflow: VLM assistance + human review of both. CulturalFrames used humans for both; its automated metrics showed weak agreement with people. This is evaluation design, not a completed scoring run.

Cinema: the architecture.
Physical features → Gemma → people.

Start with the craft inside the image
Example: compare a real frame with a copy whose grain has been removed. First measure the physical change; then use controlled image pairs to train and test the judge.
Measure image features
Lane A measures 11 descriptors: grain, colour, lighting, depth of field and other craft cues.
Controlled changes check that each measurement reacts as intended.
Train the AI judge
LoRA fine-tuning trials adapt Gemma 4 31B using real-versus-altered frame pairs and additional negative examples.
Gemma is the evaluator. The image-generating LoRA is a separate model.
Validate the judgments
People compare images blindly, with model names hidden, and rate film look and content.
Check agreement with Gemma and test for bias before ranking generators.
Measure the physical craft → teach Gemma what to look for → check it against people.
Completed: physical-feature pilot and Gemma fine-tuning trials. A reliable, human-validated generator ranking remains pending.

Cinema: what we have.
A pilot sample and a larger training library.
films in pilot
pilot reference frames
Eros titles available*
library frames*
Pilot generation: 150 prompts × 2 models × 2 seeds. Gemma: 9,600 scores; Muse: 231-score subset.
Pending: prepare the 550-title adaptation, reserve held-out films, validate the panel and schedule GPU capacity.

Example visual material: lighting, costume and atmosphere.
222.66M addressable → 8.29M processed → 6.43M captioned/exported. 3,008 GB; 77.6% yield.

Cinema: what gets judged?
One brief. Three images. Specific checks.
Shared brief: A wide night shot of a modern city street featuring a glass building, a police car with blue lights, and a blurred vehicle speeding past.

Dishoom

Qwen-Image-2512 · seed 42

Checkpoint 4250 · strength 0.5 · seed 42
Check the glass building, police car, blue lights and passing vehicle.
Compare night lighting, colour contrast, motion blur and wide-shot framing.
Rate blinded pairs. A selected example does not establish a winner.
Planned first wave: Qwen-Image-2512 base · current Eros LoRA · next Eros LoRA · i1-3B · FLUX.1 [Dev] · HiDream-I1-Full · Z-Image · Qwen-Image.

Text benchmark.
Planned for a later phase.
We are currently testing different OCR models, including Surya and Sarvam. We will complete the Text benchmark later, after comparing the readers and settling the evaluation approach.
Surya OCR · Sarvam · other OCR readers
Additional candidates include PaddleOCR / PaddleOCR-VL and Tesseract.
Test the readers
Run the same images through different models and compare how accurately they read Indian-language text.
Review the results
Check errors, difficult scripts and unclear labels with human review. Choose a consistent evaluation setup.
Complete the benchmark
Finalise the Text benchmark and publish comparable results once the model evaluation is ready.
Model comparison is in progress. The complete Text benchmark is deferred; no final model ranking is being presented.
Culture → Cinema → Text.
Choose your PDF.
Culture benchmark
1. Workflow with a worked example
2. Data, sample records & pending work
3. Example-based checks & next run
Cinema benchmark
1. Workflow with a worked example
2. Data, sample records & pending work
3. Example-based checks & next run
Text benchmark
OCR model testing & later-phase benchmark status
Download · 1 slide ↓Same content and order as the website. Detailed evidence notes · i1 model-comparison reference