Implementation Monitoring and Evaluation Division · BDPolicyLab
What has the issuing agency Implementation Monitoring and Evaluation Division (imed.gov.bd) published in the untruncated corpus, and what does its pilot document yield?
Basis: sourced records (URL, sha256 or locator per row), no model.
stored files, documents and chunks, and counts over them
Implementation Monitoring and Evaluation Division: stored documents
755
Implementation Monitoring and Evaluation Division: documents also in the office store
12
Implementation Monitoring and Evaluation Division: documents found in no other corpus
743
Implementation Monitoring and Evaluation Division: documents by upload year
Upload year
Files
Bytes
2024
Files: 724
Bytes: 4,152,528,805
2026
Files: 30
Bytes: 283,237,450
Implementation Monitoring and Evaluation Division: text-layer checks
Verdict
PDFs
Text
PDFs: 65
Legacy bangla
PDFs: 2
Implementation Monitoring and Evaluation Division: stored documents, first 50 of 755
Manifest line
Type
Bytes
Characters
Upload year
Upload month
Also in the office store
Text layer
Pages
Re-hash verified
4
Type: pdf
Implementation Monitoring and Evaluation Division: pilot documents and their printed fields
Pages
Chunks
Printed kind
Page
Line
15
Chunks: 14
Printed kind: প্রতিবেদন
Page: 1
Line:
Implementation Monitoring and Evaluation Division: pilot chunks
First page
Last page
Characters
Lines
Numeric tokens
Bangla share
1
Last page: 1
Characters: 2,306
Lines: 117
Where these figures come from
The source collections
Source
Publisher
Unit
Caveat
Government untruncated corpus harvest (stored copies, gap row A3)
Publisher: Issuing ministries, departments, divisions and boards of the Government of Bangladesh (stored copies)
Unit: one row per stored PDF document: manifest line, URL, sha256 at fetch, bytes, chars, issuing agency, portal host, upload year/month, registry entity binding, overlap with the stored files from government portals; source file is relative to the corpus root
Caveat: All 4,049 documents are full PDF publications; 241 match SHA-256 hashes of the stored files from government portals (BRN46). Agency assignment is verified against exact registry entities. sha256 verified is set only for re-hashed files (the 400 deterministic random sample and the pilot); the rest carry the harvest sha256.
Not held
Document and chunk text: Never stored in the lake or served; chunks are located by page range and pinned by the sha256 of their extracted text so a reader re-extracts them.
Documents about a person: CVs, personal data sheets, rosters and beneficiary lists are screened out of the pilot; their printed dates and memo lines would describe individuals.
Phone numbers: Counted per chunk and document at the build (phones_detected), never stored.
Publisher: Built from the rows above by the inventory pipeline
Unit: one row per issuing agency, with its partition file, sha256, file counts, and text layer counts
Caveat: Each agency's inventory partition is the resumable unit of the pipeline; a rerun rewrites only partitions whose rows changed.
200-chunk pilot: documents and chunks with printed fields across 14 agencies
Publisher: each document's stored PDF (PyMuPDF text layer)
Unit: one row per pilot document (fields with page and line) and per chunk (page range, counts, text sha256)
Caveat: Pilot documents were chosen to maximize agency diversity across text-layer candidates, chunked in consecutive page blocks up to 4,000 characters, hitting exactly 200 chunks total. Documents with personal data (CVs, data sheets, rosters, beneficiary lists, or >=2 mobile phone numbers on pages 1-2) are screened out before chunking. Phone numbers are never stored in any parquet or served payload.
Pipeline run log (measured)
Publisher: Measured pipeline runs on a 16-core machine
Unit: one row per stage run: items, bytes, seconds, workers
Caveat: Runs over files a previous run had read hit the page cache and overstate the rate; the projection uses the cold runs named in its basis.