1. User uploads PDF (10-200MB)
2. Backend creates async job, returns job_id
3. OCR service processes in background (Celery worker)
4. Frontend polls job status with progress
5. On completion, extracted schedules returned
6. User reviews and confirms schedule creation
PDF Upload
↓
Check for text layer (PyMuPDF)
↓
├── Has text: Extract directly
└── Scanned: Render to images @ 300 DPI
↓
Table detection (img2table / PaddleOCR Layout)
↓
OCR on table regions
↓
Pattern matching for maintenance data
↓
Structured schedule extraction
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Overview
Implement async PDF processing for owner's manuals, extracting maintenance schedule tables and creating maintenance_schedules entries.
Parent Issue: #12 (OCR-powered smart capture)
Priority: P3 - Owner's Manual OCR
Dependencies: OCR Service Container Setup, Core OCR API Integration
Scope
Async Processing Flow
Manual Extraction Endpoint
PDF Processing Pipeline
Pattern Matching
Mileage Intervals:
Time Intervals:
Service Types (map to maintenance subtypes):
Fluid Specifications:
Table Extraction
Directory Structure
Integration with Maintenance Feature
Acceptance Criteria
Technical Notes
Out of Scope