Basis: sourced records (URL, sha256 or locator per row), no model.
stored files, documents and chunks, and counts over them
reb.gov.bd: stored files
649
reb.gov.bd: PDF files
649
reb.gov.bd: first upload folder
2024/12
reb.gov.bd: last upload folder
2026/8
reb.gov.bd: files by type
Type
Files
Bytes
pdf
Files: 649
Bytes: 541,972,802
reb.gov.bd: files by upload folder, as printed
Upload folder
Files
2024/12
Files: 521
2026/0
Files: 9
2026/1
Files: 7
reb.gov.bd: text-layer checks
Verdict
PDFs
Text
PDFs: 6
No text
PDFs: 2
Legacy bangla
PDFs: 1
reb.gov.bd: stored files, first 50 of 649
Manifest line
Type
Bytes
Upload folder
Text layer
Pages
Re-hash verified
62,250
Type: pdf
Bytes: 46,532
Upload folder: 2024/12
reb.gov.bd: pilot documents and their printed fields
Upload folder
Pages
Chunks
Printed date
Page
Line
2024/12
Pages: 4
Chunks: 3
Printed date: ১/৭/২০১৪
reb.gov.bd: pilot chunks
First page
Last page
Characters
Lines
Numeric tokens
Bangla share
1
Last page: 1
Characters: 2,291
Lines: 237
Where these figures come from
The source collections
Source
Publisher
Unit
Caveat
Stored files from government portals
Publisher: Uploading offices of the national portal (stored copies)
Unit: one row per stored file: manifest line, URL, sha256 at fetch, bytes, office, upload folder, file type; source file is relative to the archive disk root
Caveat: upload folder is the object store's <year>/<number> folder exactly as it appears in the URL, not the document's own date; under 2026 the number runs 0 to 8, so it may be a zero-based month, which is not verified. The issuing body is the office-<slug> URL segment; its portal host is the slug with hyphens read as dots. Body and Place links are exact-domain matches only. sha256 verified is set only for re-hashed files (the random sample and the pilot); the rest carry the harvest's own sha256.
Not held
Document and chunk text: Never stored in the lake or served; chunks are located by page range and pinned by the sha256 of their extracted text so a reader re-extracts them.
Documents about a person: CVs, personal data sheets, rosters and beneficiary lists are screened out of the pilot; their printed dates and memo lines would describe individuals.
Phone numbers: Counted per chunk and document at the build (phones_detected), never stored.
A2 oracle_wave spec values: LLM-derived; used only to order pilot candidates, never served.
Publisher: Built from the rows above by the inventory pipeline
Unit: one row per uploading office, with its partition file and sha256
Caveat: Each body's inventory partition is the resumable unit of the pipeline; a rerun rewrites only partitions whose rows changed.
200-chunk pilot: documents and chunks with printed fields
Publisher: each document's stored PDF (PyMuPDF text layer)
Unit: one row per pilot document (fields with page and line) and per chunk (page range, counts, text sha256)
Caveat: Pilot documents were chosen in the priority order of the A2 oracle wave specs (LLM extractions, used as pointers only; no spec value is stored or served), one document per body, text-layer PDFs fully chunked in at most 4 chunks. Documents about a person (CVs, data sheets, rosters, beneficiary lists, or two or more mobile numbers on pages 1-2) are screened out before chunking. Kind keywords match only Unicode Bangla in logical order or English; many portal PDFs extract Bangla in visual order or legacy font encodings, so an absent kind is not evidence of none.
Pipeline run log (measured)
Publisher: Measured pipeline runs on a 16-core machine
Unit: one row per stage run: items, bytes, seconds, workers
Caveat: Runs over files a previous run had just read hit the page cache and overstate the rate; the projection uses the cold runs named in its basis.