AI Document Processing System (OCR + Data Extraction)

Please login or register as jobseeker to apply for this job.

TYPE OF WORK

Any

WAGE / SALARY

9-12

HOURS PER WEEK

25

DATE UPDATED

Mar 9, 2026

JOB OVERVIEW

Estimated Timeline: 4–6 weeks


PROJECT OVERVIEW

We are building an AI-powered document intelligence system that converts large volumes of structured and unstructured documents into a searchable analytics database.

The goal of Phase 1 is to build the core intelligence engine that can ingest documents, extract structured data using AI, store the results in a database, and generate basic analytics.

This is a backend-heavy project involving:

- OCR
- Document ingestion pipelines
- AI-based information extraction
- Database design
- Analytics-ready data structures

This project is ideal for a developer experienced with Python, AI APIs, and document processing systems.


WHAT YOU WILL BUILD (PHASE 1)

You will build a working backend system capable of:

1. Document Storage System
- Store uploaded PDF documents
- Track metadata for each document

2. Document Ingestion Pipeline
- Detect new uploaded files
- Register metadata
- Send documents through processing stages

3. OCR Processing
- Extract text from scanned PDFs
- Process documents page by page
- Recommended tool: Tesseract

4. AI Data Extraction
- Use an AI API to extract structured fields from text
- Convert unstructured document text into structured records

5. Database System
- PostgreSQL database
- Store structured extracted data
- Maintain references to original documents

6. Basic Analytics Layer
- Generate structured datasets ready for dashboards

7. Simple AI Query Interface
- Allow natural language questions against stored data


RECOMMENDED TECHNOLOGY STACK

Preferred technologies (open to suggestions):

- Python
- FastAPI or Flask
- PostgreSQL
- Tesseract OCR
- OpenAI or similar AI API
- AWS S3 or Supabase storage


EXAMPLE TASKS

Examples of tasks you may complete:

- Build a document ingestion pipeline
- Integrate OCR text extraction
- Write scripts that parse document text
- Use AI APIs to extract structured fields
- Design PostgreSQL tables for structured data
- Connect the system to a simple analytics dashboard


REQUIRED SKILLS

Strong experience with:

- Python
- API integration
- Database design
- Working with structured and unstructured data
- OCR and document processing
- AI APIs (OpenAI, Anthropic, etc.)

Bonus experience:

- FastAPI
- Supabase
- Data pipelines
- Document parsing
- Analytics dashboards


WHAT SUCCESS LOOKS LIKE

At the end of Phase 1 the system should be able to:

- Ingest documents
- Extract structured data
- Store the data in PostgreSQL
- Support analytics and AI queries


TIMELINE

Estimated build time:

4–6 weeks


HOW TO APPLY

Please include:

1. Examples of similar systems you have built
2. Your experience with OCR or document processing
3. Your experience using AI APIs
4. Your preferred development stack
5. Your hourly rate or fixed project estimate


We are looking for a developer who enjoys building data and AI systems, not just websites.

If the Phase 1 build goes well, there will be opportunities to continue working on future phases of the platform.

VIEW OTHER JOB POSTS FROM:
SHARE THIS POST
facebook linkedin