Occupational-health report generation
OCR, parse, write, review, issue — one Django desk for the lab.
A report desk for an occupational-health testing lab. Field and laboratory PDFs go in. OCR and a parser turn them into structured records. A Word report comes out. Writer, reviewer, and issuer sit on the same task. Python, Django, MySQL. Source delivered with the project.
The job
The lab already knew how to measure. What ate the week was typing: weather, job-day surveys, noise, dust, chemicals, radiation, and the rest of the original records, then assembling a report that still had to be reviewed and signed. The client needed a desk that follows China occupational-health report types — periodic testing, current-status evaluation, benzene, control-effectiveness — and does not ask a chemist to re-key a PDF.
The flow
Create a task against an employer: site, industry, unit nature, testing basis
Assign writer, reviewer, and issuer
Upload the original PDF pack
Send it to OCR, then parse into the matching record tables
Edit what the parser missed
Generate the Word report and download it
Review: pass, or return with a file and a remark
Task status is visible: new, uploaded, in parse, generating, reviewed. A parse or a report run can be watched and stopped. Errors stay on the task; they are not a silent empty document.
What gets stored
Employers and platforms. Instruments used on site. Testing basis. Then the original-data tables the report actually needs: weather; job-day investigation; materials and process; protection facilities; individual and fixed noise; dust and silica; benzene and other organics; carbon monoxide and hydrogen sulfide; microwave, UV, power-frequency and high-frequency fields; high temperature and WBGT; hand-transmitted vibration; illuminance and hood velocity; lab bench methods (chromatography, spectrophotometry, electrochemistry, atomic absorption). The catalog is the lab’s, not a generic form builder.
How it is put together
Django and SimplePro for the desk. MySQL for tasks, employers, and every original record. An OCR service reads the PDF; the application owns the parse, the Word build, and the review trail. Staff work in the admin. They do not learn a second tool for “the report part.”
Handover is the repo, the report templates, and a runbook for OCR and generation. If the lab adds another hazard or another report type, it is another table and another section in the Word build — same desk, same task.