BillSync
AI-powered inventory management built around document ingestion, structured extraction, human verification, and reliable operational data.
- AI document processing
- Inventory
- Multi-tenancy
- Backend systems
Supplier purchase bills carry the information inventory depends on, but a document is not structured data, and a machine's reading of it is not yet something a business can trust.
A pipeline that keeps the original document, the AI extraction, and human-verified business data as separate stages, so only reviewed data updates inventory and every step leaves history.
The problem
A supplier purchase bill describes stock that has arrived: who supplied it, which products it contains, and the line items an inventory system needs. But a bill is a document, not data. Before it can change inventory, the information in it has to be read, structured, and checked.
Unstructured document
Extracted information
Human verification
Trusted operational data
AI extraction can turn a document into structured fields, but what it produces is an interpretation, not a guarantee. A misread product or quantity that flows straight into stock levels becomes an inventory error, and every later decision builds on it. For data a business relies on, an unverified interpretation is not a sufficient source of truth.
The initial domain context is inventory for independent medical stores in India. The architecture is deliberately built around general document and inventory workflows rather than that one domain, so the same model can extend to other industries.
The core idea
BillSync treats three things that are easy to conflate as separate, each with its own responsibility.
Original document
Source evidence
The bill as it was received. Every later stage can be checked against it.
AI extraction
Machine-generated interpretation
Structured fields proposed by the system. Useful, but not yet trusted.
Verified business data
Human-confirmed operational truth
Data a person has reviewed, corrected where needed, and approved.
None of them silently stands in for another. The document stays available as evidence. The extraction records what the system proposed. Only verified data is treated as authoritative and allowed to change inventory. Keeping the stages apart is what makes it possible to answer, later, where a number came from and who confirmed it.
System workflow
A bill moves through seven stages, from upload to history.
Upload
The original supplier document enters the system.
Process
Processing can continue asynchronously after upload.
Extract
The system converts document information into structured fields.
Review
The extracted information is presented for human verification.
Verify
The user corrects or approves the extraction.
Update
Verified information becomes authoritative business data: products, batches, and inventory.
Audit
Important changes and inventory effects are retained as history.
Architecture
The diagram shows the development architecture BillSync is being built around. It describes the structure and direction of the system, not a production deployment.
Client
Uploads a supplier document
API / Upload
Accepts the upload and hands off processing
Message broker
Carries processing work to workers
Worker
Consumes work and processes the document
Extraction
Structured fields and snapshots
Document storage
Document files; MinIO in local development
Metadata
Bill records and lifecycle status
Human verification
Review, correct, approve
Verified business data
Products, batches, inventory
Data model
Extraction history and business state, kept apart
Redis
Part of the current infrastructure direction
The request path ends when the upload is accepted. Everything after the message broker happens in the background, on workers. Human verification sits between extraction and business data as an explicit step, not an optional one. PostgreSQL holds both the extraction history and the verified business state, in separate tables, so one never overwrites the other.
Asynchronous processing
Reading and extracting a document is potentially longer work than a user-facing request should wait on. BillSync separates the two.
Upload
Process document
Extract
Response
The upload request stays open until all processing has finished.
Upload
Accept request
Queue work
Background processing
Extraction
Verification
The request ends once the work is accepted and queued; processing continues independently.
The upload side acts as a producer: it accepts the document and publishes the processing work to a message broker. A worker consumes that work, processes the document, and persists the extraction result. RabbitMQ is the current message broker.
Producer
API / UploadPublishes processing work.
Message broker
RabbitMQHolds the work until a worker takes it.
Worker
ConsumerProcesses the document and persists the result.
Bill processing has lifecycle status handling around it, so the rest of the system can tell where a bill is in the flow without depending on the processing call itself.
AI document processing
AI output is treated as an intermediate representation: the starting point for review, never an automatic write to business data. This section describes the structure of the pipeline, not a specific model or provider.
Document
Document interpretation
Structured extraction
extraction_attemptsExtraction snapshot
extraction_snapshotsHuman review
Corrections
correctionsVerified data
Three tables carry the pipeline's history:
extraction_attemptstracks each extraction attempt separately from the bill it belongs to, so a bill's processing history is retained rather than overwritten.extraction_snapshotspreserve the structured data an extraction produced, as it was produced.correctionsrepresent changes a person makes during review, kept distinct from the machine output they change.
Data model
PostgreSQL is the database foundation, accessed through Drizzle ORM. The schema has eleven tables in five groups.
Business and identity
businesses- Tenants: each business on the platform.
users- People who use the system.
Document
bills- Uploaded supplier purchase bills.
bill_items- Line items that belong to a bill.
Extraction
extraction_attempts- Each attempt to extract a bill.
extraction_snapshots- Structured data an extraction produced.
corrections- Human changes made during review.
Inventory
products- Products the business stocks.
batches- Batches of a product.
inventory_movements- History of changes to inventory.
Audit
audit_events- Important events retained as history.
Conceptually, the groups follow the workflow. Businesses and users define who owns and acts on the data. A bill owns its line items and its extraction history, and corrections record what a person changed in that history. Verified data flows into products and batches, and changes to stock are recorded as inventory movements. Audit events cut across the groups.
Multi-tenancy
BillSync is designed as a multi-tenant SaaS. Each business is a separate tenant.
Platform
Business A
Business B
Business C
- ADMIN has administrative rights within its own business only, including managing that business's STAFF users.
- STAFF users belong to one business.
- Business users are not given access to any other business.
- SUPER_ADMIN belongs to the platform operator, not to a business. The architecture provides for platform-level, cross-business insights at that layer.
Tenancy is part of the foundation rather than added later: businesses and users are core tables, so tenant boundaries are expressed in the data model from the start.
Data integrity and auditability
BillSync does not overwrite an extraction with its corrected result. Each stage is kept as its own record.
Original document
The evidence.
Extraction attempt
extraction_attemptsAn attempt to read the document.
Extraction snapshot
extraction_snapshotsWhat the machine produced.
Human correction
correctionsWhat a person changed.
Verified business data
products · batchesWhat the business now treats as true.
Audit event / inventory movement
audit_events · inventory_movementsWhat happened, retained as history.
The system separates evidence, machine interpretation, human correction, and business state.
Preserving each stage is designed to give the system properties that a single overwritten record cannot:
- Traceability. Business data can be followed back through corrections and extractions to the original bill.
- Distinguishability. What the machine proposed and what a person changed remain separate facts.
- Retained processing history. Extraction attempts are tracked separately, so earlier attempts are not lost.
- Inventory as history. Inventory movements record how stock changed, not only where it stands now.
Development foundation
The database foundation is implemented and verified in the BillSync project.
Database foundation
- PostgreSQL with Drizzle ORM
- Eleven core tables across identity, documents, extraction, inventory, and audit
- A migration system
- Relationships between tables defined and covered by tests
- Automated tests for the database package
Local development infrastructure
Docker Compose runs the services the architecture depends on:
- PostgreSQL 16
- RabbitMQ 3.13
- Redis 7
- MinIO
This is the local development setup, not a production environment.
Engineering decisions
The design decisions the architecture is built on, and the reasoning behind each.
AI extraction is not authoritative business data.
Extraction is an interpretation of a document. Treating it as an intermediate result keeps unverified output out of inventory state.
Human verification comes before business state is trusted.
A person confirms or corrects the extraction, so authoritative data always has a human decision behind it.
Processing is asynchronous.
The upload request does not wait on potentially longer document processing, which workers handle independently of the user-facing action.
Extraction history is retained separately.
Attempts, snapshots, and corrections are separate records, so it stays clear what the system produced and what a person changed.
Multi-tenancy is part of the architecture, not an afterthought.
Businesses and users are foundation tables, so tenant boundaries are part of the data model from the start.
Audit and history are modeled explicitly.
Audit events and inventory movements are tables of their own, so important history is part of the data model itself.
Document processing is separated from inventory state.
Document and extraction tables are distinct from products, batches, and movements, and inventory is designed to change only through verified data.
Current status
BillSync is currently being developed. The database foundation is implemented and verified; the application workflow built on top of it is not complete.
Established
- Database foundation
- Core domain model
- Multi-tenant data model direction
- Document and extraction data model
- Local development infrastructure
Current direction
- Asynchronous processing through a message broker and workers
- Human verification before data becomes authoritative
- Tenant-scoped business roles and a platform layer
- Redis as part of the infrastructure
Planned next
- Complete the end-to-end processing flow
- Complete the verification workflow
- Connect extraction to verified inventory updates
- Expand operational workflows
- Continue hardening reliability and deployment
What's next
The next phase is to connect the database foundation to the complete document-processing and verification workflow, then harden the system around reliability, auditability, and operational inventory behavior.