Skip to main content
Kshitij Pal

Personal project

BillSync

AI-powered inventory management built around document ingestion, structured extraction, human verification, and reliable operational data.

Status
Currently being developed
Focus
  • AI document processing
  • Inventory
  • Multi-tenancy
  • Backend systems

Problem

Supplier purchase bills carry the information inventory depends on, but a document is not structured data, and a machine's reading of it is not yet something a business can trust.

Solution

A pipeline that keeps the original document, the AI extraction, and human-verified business data as separate stages, so only reviewed data updates inventory and every step leaves history.

01

The problem

A supplier purchase bill describes stock that has arrived: who supplied it, which products it contains, and the line items an inventory system needs. But a bill is a document, not data. Before it can change inventory, the information in it has to be read, structured, and checked.

From document to operational data
  1. Unstructured document

  2. Extracted information

  3. Human verification

  4. Trusted operational data

AI extraction can turn a document into structured fields, but what it produces is an interpretation, not a guarantee. A misread product or quantity that flows straight into stock levels becomes an inventory error, and every later decision builds on it. For data a business relies on, an unverified interpretation is not a sufficient source of truth.

The initial domain context is inventory for independent medical stores in India. The architecture is deliberately built around general document and inventory workflows rather than that one domain, so the same model can extend to other industries.

02

The core idea

BillSync treats three things that are easy to conflate as separate, each with its own responsibility.

Three distinct representations
  • Original document

    Source evidence

    The bill as it was received. Every later stage can be checked against it.

  • AI extraction

    Machine-generated interpretation

    Structured fields proposed by the system. Useful, but not yet trusted.

  • Verified business data

    Human-confirmed operational truth

    Data a person has reviewed, corrected where needed, and approved.

None of them silently stands in for another. The document stays available as evidence. The extraction records what the system proposed. Only verified data is treated as authoritative and allowed to change inventory. Keeping the stages apart is what makes it possible to answer, later, where a number came from and who confirmed it.

03

System workflow

A bill moves through seven stages, from upload to history.

Core workflow
  1. Upload

    The original supplier document enters the system.

  2. Process

    Processing can continue asynchronously after upload.

  3. Extract

    The system converts document information into structured fields.

  4. Review

    The extracted information is presented for human verification.

  5. Verify

    The user corrects or approves the extraction.

  6. Update

    Verified information becomes authoritative business data: products, batches, and inventory.

  7. Audit

    Important changes and inventory effects are retained as history.

04

Architecture

The diagram shows the development architecture BillSync is being built around. It describes the structure and direction of the system, not a production deployment.

System architecture
Development architecture
    • Client

      Uploads a supplier document

    • API / Upload

      Accepts the upload and hands off processing

    • Message broker

      RabbitMQ

      Carries processing work to workers

    • Worker

      Consumes work and processes the document

    • Extraction

      Structured fields and snapshots

    • Document storage

      Object storage

      Document files; MinIO in local development

    • Metadata

      Bill records and lifecycle status

    • Human verification

      Review, correct, approve

    • Verified business data

      Products, batches, inventory

    • Data model

      PostgreSQL · Drizzle ORM

      Extraction history and business state, kept apart

  • Redis

    Part of the current infrastructure direction

The request path ends when the upload is accepted. Everything after the message broker happens in the background, on workers. Human verification sits between extraction and business data as an explicit step, not an optional one. PostgreSQL holds both the extraction history and the verified business state, in separate tables, so one never overwrites the other.

05

Asynchronous processing

Reading and extracting a document is potentially longer work than a user-facing request should wait on. BillSync separates the two.

Synchronous
  1. Upload

  2. Process document

  3. Extract

  4. Response

The upload request stays open until all processing has finished.

Asynchronous
  1. Upload

  2. Accept request

  3. Queue work

  4. Background processing

  5. Extraction

  6. Verification

The request ends once the work is accepted and queued; processing continues independently.

The upload side acts as a producer: it accepts the document and publishes the processing work to a message broker. A worker consumes that work, processes the document, and persists the extraction result. RabbitMQ is the current message broker.

Producer, broker, worker
  1. ProducerAPI / Upload

    Publishes processing work.

  2. Message brokerRabbitMQ

    Holds the work until a worker takes it.

  3. WorkerConsumer

    Processes the document and persists the result.

Bill processing has lifecycle status handling around it, so the rest of the system can tell where a bill is in the flow without depending on the processing call itself.

06

AI document processing

AI output is treated as an intermediate representation: the starting point for review, never an automatic write to business data. This section describes the structure of the pipeline, not a specific model or provider.

Extraction pipeline
  1. Document

  2. Document interpretation

  3. Structured extractionextraction_attempts

  4. Extraction snapshotextraction_snapshots

  5. Human review

  6. Correctionscorrections

  7. Verified data

Three tables carry the pipeline's history:

  • extraction_attempts tracks each extraction attempt separately from the bill it belongs to, so a bill's processing history is retained rather than overwritten.
  • extraction_snapshots preserve the structured data an extraction produced, as it was produced.
  • corrections represent changes a person makes during review, kept distinct from the machine output they change.

07

Data model

PostgreSQL is the database foundation, accessed through Drizzle ORM. The schema has eleven tables in five groups.

  • Business and identity

    businesses
    Tenants: each business on the platform.
    users
    People who use the system.
  • Document

    bills
    Uploaded supplier purchase bills.
    bill_items
    Line items that belong to a bill.
  • Extraction

    extraction_attempts
    Each attempt to extract a bill.
    extraction_snapshots
    Structured data an extraction produced.
    corrections
    Human changes made during review.
  • Inventory

    products
    Products the business stocks.
    batches
    Batches of a product.
    inventory_movements
    History of changes to inventory.
  • Audit

    audit_events
    Important events retained as history.

Conceptually, the groups follow the workflow. Businesses and users define who owns and acts on the data. A bill owns its line items and its extraction history, and corrections record what a person changed in that history. Verified data flows into products and batches, and changes to stock are recorded as inventory movements. Audit events cut across the groups.

08

Multi-tenancy

BillSync is designed as a multi-tenant SaaS. Each business is a separate tenant.

Tenant structure
  • PlatformSUPER_ADMIN

    • Business A

      • ADMIN
      • STAFF
    • Business B

      • ADMIN
      • STAFF
    • Business C

      • ADMIN
      • STAFF
  • ADMIN has administrative rights within its own business only, including managing that business's STAFF users.
  • STAFF users belong to one business.
  • Business users are not given access to any other business.
  • SUPER_ADMIN belongs to the platform operator, not to a business. The architecture provides for platform-level, cross-business insights at that layer.

Tenancy is part of the foundation rather than added later: businesses and users are core tables, so tenant boundaries are expressed in the data model from the start.

09

Data integrity and auditability

BillSync does not overwrite an extraction with its corrected result. Each stage is kept as its own record.

From evidence to business state
  1. Original document

    The evidence.

  2. Extraction attemptextraction_attempts

    An attempt to read the document.

  3. Extraction snapshotextraction_snapshots

    What the machine produced.

  4. Human correctioncorrections

    What a person changed.

  5. Verified business dataproducts · batches

    What the business now treats as true.

  6. Audit event / inventory movementaudit_events · inventory_movements

    What happened, retained as history.

The system separates evidence, machine interpretation, human correction, and business state.

Preserving each stage is designed to give the system properties that a single overwritten record cannot:

  • Traceability. Business data can be followed back through corrections and extractions to the original bill.
  • Distinguishability. What the machine proposed and what a person changed remain separate facts.
  • Retained processing history. Extraction attempts are tracked separately, so earlier attempts are not lost.
  • Inventory as history. Inventory movements record how stock changed, not only where it stands now.

10

Development foundation

The database foundation is implemented and verified in the BillSync project.

Database foundation

  • PostgreSQL with Drizzle ORM
  • Eleven core tables across identity, documents, extraction, inventory, and audit
  • A migration system
  • Relationships between tables defined and covered by tests
  • Automated tests for the database package

Local development infrastructure

Docker Compose runs the services the architecture depends on:

  • PostgreSQL 16
  • RabbitMQ 3.13
  • Redis 7
  • MinIO

This is the local development setup, not a production environment.

11

Engineering decisions

The design decisions the architecture is built on, and the reasoning behind each.

  1. AI extraction is not authoritative business data.

    WhyExtraction is an interpretation of a document. Treating it as an intermediate result keeps unverified output out of inventory state.

  2. Human verification comes before business state is trusted.

    WhyA person confirms or corrects the extraction, so authoritative data always has a human decision behind it.

  3. Processing is asynchronous.

    WhyThe upload request does not wait on potentially longer document processing, which workers handle independently of the user-facing action.

  4. Extraction history is retained separately.

    WhyAttempts, snapshots, and corrections are separate records, so it stays clear what the system produced and what a person changed.

  5. Multi-tenancy is part of the architecture, not an afterthought.

    WhyBusinesses and users are foundation tables, so tenant boundaries are part of the data model from the start.

  6. Audit and history are modeled explicitly.

    WhyAudit events and inventory movements are tables of their own, so important history is part of the data model itself.

  7. Document processing is separated from inventory state.

    WhyDocument and extraction tables are distinct from products, batches, and movements, and inventory is designed to change only through verified data.

12

Current status

BillSync is currently being developed. The database foundation is implemented and verified; the application workflow built on top of it is not complete.

  • Established

    • Database foundation
    • Core domain model
    • Multi-tenant data model direction
    • Document and extraction data model
    • Local development infrastructure
  • Current direction

    • Asynchronous processing through a message broker and workers
    • Human verification before data becomes authoritative
    • Tenant-scoped business roles and a platform layer
    • Redis as part of the infrastructure
  • Planned next

    • Complete the end-to-end processing flow
    • Complete the verification workflow
    • Connect extraction to verified inventory updates
    • Expand operational workflows
    • Continue hardening reliability and deployment

13

What's next

The next phase is to connect the database foundation to the complete document-processing and verification workflow, then harden the system around reliability, auditability, and operational inventory behavior.