# Eros IndicGen - founder briefing evidence notes

Snapshot inspected: 15 September 2026. This briefing describes the source and saved artifacts accessible in the three benchmark checkouts on dev 2. It does not claim to have rerun GPU experiments, a human panel or the full test suites.

## Repository provenance

Remote root: `/home/thilak.arjun/IndicGen-Benchmark/benchmark/`.

| Benchmark | Directory | Inspected revision |
|---|---|---|
| Cultural | indic_cultural_bench | outer repository 4de6678; workspace files inspected directly |
| Text | indic_text_bench | 5fc092c, indictext |
| Cinema | indic_cinema_bench | 4e70827, rating-panel-v12 |

The original repositories and their nested Git configuration were not modified. The briefing is a separate website.

## Cultural: current measured inventory

- `data/prompts/v0.3/canonical.jsonl`: 237 records, 237 scene groups, all marked draft; 34 distinct state/UT labels; 10 categories.
- Categories: weddings 24, dance 26, geography 21, urban 23, attire 23, food 24, rural 27, architecture 26, festivals 22, mythology 21.
- `data/prompts/v0.3/translations.jsonl`: 1,356 translation/mode rows, all marked needs_review. This is not a count of scenes or approved translations.
- `data/verify/v0.3/verification.jsonl`: 198 records, all matching current canonical source hashes. Status counts: 143 machine_verified, 55 needs_expert. The remaining 39 of 237 canonical scenes have no entry in this ledger. All 198 records identify `google/gemma-4-31B-it`.
- `verify_report.json` is the last batch summary (56 verified, 22 expert, 120 skipped), not the whole-corpus status. The briefing uses the full verification ledger.
- Configuration target: 500 scenes; five-language design (English, Hindi, Tamil, Telugu, Bengali). Targets are not completed outputs.
- Source registry: UNESCO ICH, Wikipedia, Wikidata, Wikimedia Commons, Sahapedia, IGNCA, ASI, Census of India, Ministry of Tribal Affairs, Cultural Atlas. Some are cite-only sources. The Atlas is an optional local cache and contextual observance inventory, not sufficient regional visual evidence by itself.
- Authoring uses Claude Code with no globally pinned author model. Some translation provenance records identify `claude-fable-5-1`; that does not establish one model across the corpus.
- Ten ingest gates check structure, grounding, sources, script, taxonomy, duplication, modes and balance. Gemma checks explicit claims against evidence. Reviewers settle cultural and linguistic questions and blind identification. The audit does not override people.
- No image generation or image scoring is implemented in this repository. Planned image scoring is 30% semantic + 30% cultural + 30% core attribute coverage + 10% visual quality, on explicit rows only, scaled to 0-100. Other modes diagnose defaults and inference. No completed image-model ranking is claimed.

Primary files: README.md; configs/benchmark.yaml; data/reference/sources.yaml; docs/verification.md; docs/evaluation.md; canonical, translation and verification files above.

## Text: saved v2 OCR baseline

Primary files: `experiments/indictext_v2/{config.yaml,sample_mix.json,metrics.json,report.md,generations/manifest.jsonl}`; `data/prompts/indictext_v2.jsonl`; `docs/scoring.md`; `docs/prompt-recovery-2026-09-11.md`; `configs/evaluators.yaml`.

- 2,000 source images, seed 42, 12 languages. Prompt and manifest files each have 2,000 records. Manifest entries describe imported Hugging Face data; the model label `indictext_v2` is an experiment/dataset label, not evidence of an image generator.
- Source counts: `dhruvil237/indicbench` 407; `darknight054/indic-mozhi-ocr` 407; `Bhashini-IITJ/BharatSceneTextDataset` 407; `darknight054/indicstr12-crops` 407; `darknight054/indicstr12-full` 372.
- Headline OCR: Tesseract 5. Saved score 0.3301590704 = 33.0%; character accuracy 0.3946864206 = 39.5%; word accuracy 0.2656317203 = 26.6%; exact match 0.14 = 14.0%.
- The score equally weights mean grapheme-character accuracy and mean word accuracy. Raw error rates can exceed 100% when insertions dominate; each sample's accuracy is clipped. Thus 1 minus the average error rate is not the reported average accuracy.
- Language scores and counts: Marathi 48.0% (167); Hindi 47.0% (159); Bengali 38.2% (179); Kannada 34.4% (174); Punjabi 34.2% (190); Assamese 32.3% (170); Gujarati 32.2% (168); Tamil 29.9% (179); Telugu 28.1% (179); Malayalam 24.9% (167); Odia 24.2% (164); Urdu 17.4% (104).
- PaddleOCR/PaddleOCR-VL and Surya are diagnostic OCR implementations. Qwen2.5-VL (3B/7B/72B options), Gemma 4 31B, Muse Glimmer 30B and GPT-4o appear as configured VLM adapters. Configuration is not proof of a completed comparable run.
- V1's saved 30.9% result has a different sample composition/coverage and is not shown as a controlled 2.1-point improvement in v2.
- Full v2 image collections and raw OCR are not bundled in the inspected clone. Selected v2 source-image paths were unavailable. Historical online OCR and manual input examples are labelled separately in the presentation. A fresh clone can show saved charts without being a complete reproduction bundle.
- Mock demo scores and the incomplete historical full-evaluation dataset are not used as founder results.

## Cinema: pilot evidence and current readiness

Primary files: `docs/pilot_69/{report.json,dashboard.html,gate_result.json,source_guard_verdict.json}`; `docs/pilot_69/lane_b/{lane_b_run_summary.json,bakeoff.json}`; `configs/lane_b_templates.yaml`; `docs/RATING_UI.md`; `src/indic_cinema_bench/preprocessing/gemma_box3.py`.

- Completed pilot: 69 film titles x 100 sampled frames = 6,900 real frames. Sampling is deterministic and scene-stratified.
- Pilot preprocessing normalises full frames and extracts subject, background and highlight crops. Five controlled craft alterations plus an identity re-encode test the descriptors.
- Lane A contains 11 descriptors. The hard normalisation gate passed without bypass. Byte-length-including diagnostics failed. A hard-gate pass is not a blanket claim that every diagnostic or selectivity test passed.
- Lane B pilot: 150 caption-derived prompts; Qwen-Image-2512 base and Indic-cinematic LoRA checkpoint 4250 at scale 0.5; two seeds. Pilot judging: Gemma 4 31B, 9,600 scores over 2,400 pairs and four dimensions; Muse Glimmer 30B, only 231 scores. These counts have different denominators and should not be presented as equal-size judge comparisons.
- Gemma original-over-ablation accuracy: 0.66 over 6,000 scores, 95% interval approximately 64.8-67.2%. It is a diagnostic comparison with deliberately changed controls, not a validated generator leaderboard.
- F = filmicness; I = Indian material authenticity; S = style fidelity; D = diegesis (a moment occurring within a scene, rather than a posed or poster-like image). D does not mean depth.
- Current template set v4 adds P = prompt adherence, eligible only on generated-vs-generated comparisons. Do not mix this current five-dimension protocol with historic four-dimension pilot scores.
- Current rating code derives human wording from the judge templates and uses the same subject crop resolver, with matched full-frame-plus-subject views. The latest code includes Gemma box3 SubjectLocator crop functionality; its existence is not a new completed accepted benchmark result.
- Current template comments and rating documentation identify earlier v5-v7 human-labelled material as rating-UI development/test output. Historical human preference, agreement and acceptance numbers that relied on it are excluded from this briefing. A real blinded panel is needed for acceptance.
- Earlier judge status narratives include withdrawn or superseded conclusions. The briefing uses saved pilot measurements and current code-level caveats instead of presenting old acceptance claims as current truth.

## Images and branding

- Eros Innovation logo: official website https://erosinnovation.com/ ; official header asset https://erosinnovation.com/wp-content/uploads/2025/01/eros_innovation_logo.webp (1000 x 393). Replaced the previous brand asset at user request.
- Cultural photograph: Kathakali performance in Kochi, Kerala, by Arian Zwegers. Source https://commons.wikimedia.org/wiki/File:Kochi,_Kathakali_performance_(6317474153).jpg . CC BY 2.0: https://creativecommons.org/licenses/by/2.0/ . Used as an illustrative cultural reference, not a generated or scored output. The photo is not asserted to depict a specific prompt record.
- Cinema images and altered variants: extracted from embedded images in the inspected repository's saved `docs/pilot_69/dashboard.html`. Cover frame: Devdas. Slide 11 uses a real Bajirao Mastani reference from the separate model-development archive to illustrate five craft cues; it is not claimed as a scored pilot ablation. Source: /home/thilak.arjun/qwen2512-indic-cinematic-v1/benchmark/site/samples/Bajirao_Mastani__1023354__0714be81ccbf__real.jpg. Slide 12 opens with the Phuntroo real/base/LoRA portrait comparison from the pilot dashboard (gallery index 0). The alternate selectable comparison is Son of Sardaar (gallery index 4). Model outputs and seeds retain the archive labels. Display examples are low-resolution archive previews.
- Tamil/Hindi sign images: repository `experiments/manual_image_ingestion/IT_ta_01_0000_s0.jpg` and `IT_hi_01_0000_s0.png`. These are input illustrations with no claim about their production model or inclusion in scored v2.
- Urdu crop: repository online OCR experiment image `IT_ur_01_9ff04b8194ce5fbab556d39da5faeab4_s0.jpg`, distinct from the v2 results.

## Delivery checks

The website has desktop/mobile layouts, keyboard presentation navigation, reduced-motion support, cultural mode tabs, five explained craft cues and selectable real/base/LoRA examples. The PDF is a 15-page landscape rendering of the same narrative with default examples. Browser checks cover image loading, JavaScript errors, interactive controls and mobile width. PDF pages were rendered and visually inspected after pagination fixes.
# Expansion: pilot evidence, Eros data and full-run plan — 15 September 2026

## Three distinct evidence scopes

The three benchmark checkouts describe development/pilot work, not a completed multi-model leaderboard. The cultural corpus is still being reviewed; text v2 is an OCR evaluation on existing source images; the cinema 69-film run tests descriptors and judge behaviour. The new generation programme is proposed work.

A separate model-development workspace, `/home/thilak.arjun/qwen2512-indic-cinematic-v1/benchmark/`, contains a saved LoRA study. Its README and `runs/full/scorecard.md` report CMMD (lower is closer to real film) of 1.2168 for base, 1.3302 for the stronger prompt-engineering arm, and 0.5088 for checkpoint 4250 at scale 0.50. It uses 1,050 prompts from seven held-out films. These are saved study results, not a new run for this presentation.

`runs/full/human_results.json` reports film-look preference of 65.8537% against base (164 comparisons; 95% interval 58.3304–72.5860%) and 64.7436% against stronger prompt engineering (156 comparisons; interval 55.1851–72.4025%). The README reports 17 participating raters and 320 screened film-look comparisons. This is preliminary: the identical-pair control selected left 35.4167% of non-tie votes and its `ok` flag is false. Caption matching is not significantly improved. The deck does not claim production acceptance. This saved study is separate from the rating-UI test labels excluded from the newer cinema benchmark.

## Team-supplied programme and operations snapshot

The user supplied the following programme figures on 15 September 2026: a 100-title experiment; 550 quality-cleared Eros titles available for LoRA training; 1.3 million frames from that library; a training pipeline proven on a 3-million-image workload. These are reported programme figures, not independently recounted assets. The inspected workspace describes a 69-film run; the 100-title programme statement must not be silently treated as its exact training manifest. 550/100 = 5.5 times the available titles, not a measured model-quality increase. No completed 550-title LoRA is asserted.

Wider corpus counters supplied: addressable pool 222,663,181; processed 8,288,862; captioned/exported 6,434,437; export 3,008 GB; 9,594 completed work units across 533 cycles; 61 failed; remaining 214,342,656; reported progress 3.74%, running. Yield rounds to 77.6%; failure percentage rounds to 0.6%. These are a supplied snapshot, not a live telemetry feed.

Counter reconciliation: identified minus processed equals 214,374,319, which is 31,663 higher than the supplied remaining counter. Processed/identified is about 3.72%, rather than the reported 3.74%. The deck preserves the supplied progress and remaining numbers as explicitly reported counters. Their stage definitions or capture times require reconciliation before treating them as a single audited accounting identity. Addressable data is not synonymous with quality-approved or training-ready data. The wider corpus is separate from the Eros movie library; no claim that all corpus images are owned Eros frames is made.

GPU availability is the team's reported current scaling constraint. A low failure rate is an encouraging reliability signal, not sufficient proof of caption accuracy, end-to-end data quality, or unattended operational safety. Better model quality is the aim to test. A proprietary foundation model is a longer-term ambition requiring a costed training programme and validated gains, not a guaranteed consequence of more GPU hours.

## i1 reference and proposed comparison roster

Primary source: [i1: A Simple and Fully Open Recipe for Strong Text-to-Image Models](https://arxiv.org/html/2606.11289v1#S6.T8), Princeton, June 2026, Table 8. Its 17 model rows (including i1) are: GPT Image 1 [High], Seedream 3.0, FLUX.1 [Dev], SD3 Medium, Janus-Pro-7B, BAGEL, HiDream-I1-Full, Lumina-Image 2.0, Z-Image, Qwen-Image, BLIP3o-4B, PixNerd, DeCo, BLIP3o-N-S, BLIP3o-N-G-G, BLIP3o-N-G-T, and i1. The paper table mixes author-measured and previously reported scores; none are presented as Eros benchmark results. i1 is distinct from HiDream-I1.

The presentation proposes those 17 external entrants plus Qwen-Image-2512 base, the current Eros LoRA and a future LoRA using the 550-title pool: 20 entrants. First wave: the three local entrants, i1, FLUX.1 [Dev], HiDream-I1-Full, Z-Image and Qwen-Image. Remaining 12 entrants are an expansion stage. Model/API access, terms, exact revisions and compute capacity must be resolved before execution. This is a proposed comparison plan, not a claim of downloaded models, completed runs, or approved spending.

Planning arithmetic: 500 reviewed cultural scenes × 5 languages = 2,500 explicit prompts; 2,000 newly authored/balanced text-generation prompts across 12 languages; 1,000 frozen held-out cinema prompts. Total 5,500 prompts × 2 outputs × 8 entrants = 88,000 generated images, or ×20 entrants = 220,000. The culture 500-scene target already exists. The text and cinema generation counts and staging are proposed planning assumptions. Additional prompt modes, ablations, controls, reruns, source images and human judgments are excluded from this generated-image count. Test films must be excluded from training. Common seeds apply where supported; freeze versions and document model-specific settings. No GPU-hour or cost estimate is implied.

---
# Worked-example edition — nine slides, three per benchmark

**Named cultural input sources:** Culture slide 1 names Wikipedia, Wikidata, Cultural Atlas, UNESCO, Wikimedia Commons, Sahapedia, IGNCA, ASI, Census of India and Ministry of Tribal Affairs from the inspected reference registry. Indian tourism references were explicitly supplied by the team for this briefing; no tourism URL was present in the inspected v0.3 canonical sources or registry, so their exact pages and ingestion status have not been independently verified here. This is a programme-level source list, not a claim that every source supports every scene. The displayed Magh Bihu record cites its specific article. Cultural Atlas supplies broad context and cannot alone substantiate regional visual details.

**Current Text scope correction:** The team clarified that Text evaluates OCR readers, including Sarvam and Surya OCR alongside other models. All three current Text slides now describe images plus correct reference text → OCR readers → scoring → human review → model comparison. The saved 33.0% / 39.5% / 26.6% figures remain Tesseract 5 results; they are not attributed to Sarvam or Surya. No particular Sarvam product version or completed Sarvam/Surya score has been verified in this presentation audit. Earlier text-generation plans, generated-image counts and corresponding aggregate run envelopes below are historical planning notes, superseded for the current Text section.

Cinema slide 1 now explicitly presents physical measurement (Lane A) → Gemma judge fine-tuning (Lane B) → human evaluation. The physical pilot measured 11 descriptors; Gemma 4 31B LoRA trials used controlled frame pairs and additional negatives. `JUDGE_STATUS_2026-09-06.md` records v3–v6 trials, including an abandoned run. These establish that fine-tuning was attempted, not that an accepted judge resulted. Later documentation supersedes that older note's human-label and negative-space conclusions; no such older preference or acceptance figures are reused here. Gemma is an evaluator and is distinct from the cinematic image-generator LoRA. Human acceptance remains pending under the current rating documentation.

Culture slide 3 now explains explicit, caption, implicit and bare prompts using illustrative variants of the Magh Bihu scene. Explicit supplies cultural details; caption supplies the cultural name; implicit removes the festival and object names while keeping contextual clues; bare asks for a generic Indian harvest festival. These variants illustrate the mode design and have not been released or scored. Human experts must check recognisability of implicit prompts. Caption, implicit and bare are diagnostics and do not inherit all explicit-prompt scoring obligations.

The current presentation follows Culture → Cinema → Text. Each topic has three slides: architecture with a worked example; data and sample material; example-based checks and the proposed next run. Separate topic PDFs contain the same slide content as the nine-page combined PDF.

Culture uses actual canonical draft `IG_G_001139`, a Magh Bihu scene near Nagaon. Its required core details are a bamboo-and-hay meji bonfire, a bhelaghar hut beside it, and women offering til pitha and coconut laru. The displayed prompt is shortened. The record is machine-verified but remains a draft. The generic-bonfire failure described on slide three is hypothetical, not observed generated output. Human approval and cultural generator evaluation are still pending.

The bhelaghar reference photograph is [Magh bihu 04.jpg by Abhi179](https://commons.wikimedia.org/wiki/File:Magh_bihu_04.jpg), [CC BY-SA 4.0](https://creativecommons.org/licenses/by-sa/4.0/). It is displayed with a crop; the downloaded asset remains the original image. The photograph and displayed crop retain that license. It illustrates a visual requirement, rather than an image produced or scored by the benchmark.

Cinema uses the actual Phuntroo reference and archived Qwen-Image-2512 base / checkpoint-4250 LoRA comparison, both displayed generated examples at seed 42. Individual showcase examples establish no ranking. The Bajirao Mastani reference belongs to the model-development archive, not the scored 69-film pilot ablation set.

The Hindi and Tamil examples come from the separate repository input set and are not claimed as scored v2 images. The Hindi matching and missing-vowel examples are illustrative scoring cases, not real OCR predictions. The numerical OCR baseline still refers to Tesseract 5 on the frozen v2 source images.

The following detailed programme notes retain the underlying evidence and planning assumptions, including figures omitted from the shorter slides.

---


## Cultural expectation scoring correction (15 September 2026)

Culture slide 3 now explains explicit and implicit expectations within the same image-prompt evaluation. It replaces the earlier four-prompt-mode explanation, which did not address the requested scoring distinction.

Source: Nayak et al., *CulturalFrames: Assessing Cultural Expectation Alignment in Text-to-Image Models and Evaluation Metrics*, arXiv:2506.08835v1, sections 3.3 and 6. https://arxiv.org/html/2506.08835v1

The paper's human alignment scale is 0 / 0.5 / 1. For imperfect alignment, raters identify explicit, implicit, or both kinds of failure and explain their judgment. These are failure labels attached to an alignment rating, not two separately reported numerical scores. Its VLM-based metrics do not replace human cultural assessment.

The VLM-assisted explicit checklist and human review division shown for Eros is a proposed workflow, not a claim that the paper used this division or that our benchmark has completed it. Humans must validate explicit checks as well as implicit judgments. The shortened Magh Bihu prompt is an illustration adapted from IG_G_001139, not a scored result or a CulturalFrames prompt. Evidence and regional variation determine which contextual expectations are appropriate for the scene.


## Text benchmark status update (15 September 2026)

The founder presentation now has seven slides: Culture 3, Cinema 3 and Text 1. The single Text slide supersedes the earlier three-slide Text presentation. Per the team update, Surya, Sarvam and other OCR readers are being tested; completion of the Text benchmark is deferred until model comparison and the evaluation approach are ready. Candidate readers are not claimed to have completed comparable runs. Historical baseline figures remain background evidence, not a final ranking.


## Cinema example replacement (15 September 2026)

Cinema slide 3 now uses the saved Dishoom night-street triplet from docs/pilot_69/dashboard.html, laneb_gallery index 2, prompt p0060_8e2a9cd1f52e, movie Dishoom__1054611. The gallery explicitly labels generated images as seed 42. Its entrants identify Qwen-Image-2512 base and cinematic adapter checkpoint 4250 at LoRA scale 0.5. The shared caption is shown on the slide, with checks updated to buildings, police lights, motion blur and framing. These are unchanged saved images, chosen as a clearer presentation example; the selection does not establish comparative performance.


## Restored Assamese clothing reference (15 September 2026)

Culture slide 2 restores the earlier photograph of Bihu dancers from Dhakuakhana, Assam, by Nayan j Nath, replacing the bhelaghar image. Source: https://commons.wikimedia.org/wiki/File:Bihu_dancers_from_Dhakuakhana_Assam_India_2.jpg . License: CC BY-SA 4.0, https://creativecommons.org/licenses/by-sa/4.0/ . The original image bytes are restored from the presentation history, displayed with a crop, and remain under that license. This is a regional clothing illustration, not documentation of the specific Magh Bihu ritual or an evaluated model output.
