PAWAN
I build data platforms and the governance that lets teams use AI agents with real databases: on-prem pipelines, a semantic layer that compiles approved queries, and tamper-evident audit records. These open-source tools are used by other developers, with thousands of package installs and outside contributors. Based in Greater Noida, Delhi NCR. Compliance shouldn't mean spreadsheets, and AI shouldn't require the cloud.
About Me
MSc Data Analytics · Open-source maintainer
I build governed data and AI systems, including open-source tools for SQL access, privacy controls, agent audit and manufacturing analytics. These projects are built in the open and designed to run on-prem.
// SHIPPED
- sql-steward — a semantic-layer MCP server where an AI agent queries a real database and never writes SQL. On PyPI and in the official MCP registry.
- Its on-prem governance layer: schema-scout maps, query-warden gates access, pii-veil redacts, agent-blackbox audits. Seven packages on PyPI in all.
- The governed-agent-stack: FloorMind, elenchus and the audit layer in one monorepo behind a single Next.js console, every backend proxied server-side.
- ETL from ERP systems into dbt marts with SLO monitoring
// ARCHITECTURE
- LangGraph agents over MCP servers; ChromaDB and pgvector RAG; Ollama models
- Schema-aware validation gate with a self-correcting repair loop
- Every answer records its model and prompt — reproducible, audited
- Fully on-prem; no batch data leaves the building
// UPSTREAM
- 36 merged in drt (three releases, destinations, parallel orchestration, CI), plus merged fixes upstream: Apache Superset, sqlglot, SQLFluff, scanapi, fpdf2, DuckDB, dlt, Open Food Facts, Electricity Maps
- In review: SQLMesh, sops, Octopus Energy, dlt, textual-fastdatatable
- Read the source, find the gap, ship the fix
Work History
The paid work behind the open-source one — the domain problems that shaped the tools.
Data Engineer
- End-to-end pipelines from source systems into governed warehouses — SQL Server, Postgres, BigQuery, DuckDB.
- Manufacturing compliance dashboards: BRC/HACCP, NL query for auditors, /metrics Golden Signals, PDF audit reports.
- UK Crime Pipeline: Police UK API to Postgres and BigQuery, 6 dbt marts, 65 tests, Polars-based alternative ingestion.
Academic Background
Data analytics formal study plus the certifications that mattered for the day job.
MSc Data Analytics — Aston University
Coursework across statistical learning, big-data processing (Spark), and database systems, with a research dissertation applying machine learning (Random Forest) to the effects of ethnicity on student behaviour and outcomes.
BCA — Artificial Intelligence and Machine Learning
Algorithms, OOP, RDBMS, software engineering. Final-year project: a client-server inventory tracker with a JSP front and MySQL back.
Microsoft Certified: Power BI Data Analyst Associate (PL-300)
Modelling, DAX, Power Query M, time intelligence, Import vs DirectQuery. Used daily for the manufacturing reporting layer.
Other certifications
Google Data Analytics Professional Certificate · Azure Data Engineering (DP-203 path) · AWS Cloud Practitioner.
Projects
What I've shipped, and the adoption it has picked up.
Drive it. Ask a metric and the question is compiled to a bounded query, run, and logged. Request PII and the guard refuses it at compile time, before a row is ever read. See the full ungoverned-vs-governed side-by-side → · Try the policy live: type SQL, watch it get refused →
sql-steward
The core product: a governed MCP server where an AI agent queries a real database and never writes SQL. Every query is compiled from a semantic layer you control — entities, joins, metrics, PII tags — so there is no run_sql tool to misuse. Blocked PII is refused before the query runs, declared data-quality checks gate the results, and init --from-db drafts a reviewable layer straight from a live schema, so it scales past a toy database. Same tools across SQL Server, Postgres, and SQLite. Listed in the official MCP registry.
Four small, on-prem tools that plug into sql-steward — each does one job, each ships on PyPI, none needs the cloud.
schema-scout map
Reverse-engineers a 150+ table SQL Server database into an AI-ready catalog, inferring foreign keys and flagging PII. This is what init --from-db reads.
query-warden access
YAML role-based access control for SQL, enforced before the query runs. A role can be denied a table or a single column and never sees a row of it.
pii-veil redact
Column-level PII masking and refusal. Tagged columns are masked in the result or blocked outright at compile time, before a query ever touches the data.
agent-blackbox audit
Append-only, hash-chained ledger: each row stores the hash of the one before it, so any later edit breaks the chain and verify() points at it. One SQLite file, zero deps.
The same argument as the governance layer, moved to money: an agent should not be trusted with a payment rail just because it can reach one. Three tools, each public.
steward delegate
An agent that spends someone else's money. It starts from the asymmetry the socks-buying demos skip: the person spending and the person paying are different people, and the payer cannot approve every purchase. The spender texts; the sponsor's policy decides. The sponsor sees decisions, the ledger and escalations — never the conversation.
pay-warden gate
The policy engine underneath, as an MCP server. Budgets, per-purchase caps, merchant deny-lists, velocity windows and a human-approval threshold are evaluated before a request reaches the payment rail. Amounts are exact decimals converted through declared rates, so pricing in a weaker currency cannot slip a cap.
payoptimize route
Once a purchase is allowed, it still has to authorise. Routes each payment to the provider most likely to accept it — discounted Thompson sampling over live authorisation rates, with a decline-aware cascade that retries the next-best rail instead of failing the customer.
UK Crime Pipeline
End-to-end data pipeline. Police UK API to PostgreSQL and BigQuery. 6 dbt marts (outcome analysis, YoY trends), 65 tests. Polars-based alternative ingestion. Declarative validation, SLO monitoring, pipeline maturity scorecard.
Manufacturing Compliance Dashboard
BRC/HACCP food safety compliance. MCP server exposes 5 compliance tools for LLM agents. NL query interface for auditors. SLO monitoring (temp 95%, traceability 90%), z-score anomaly detection, PDF audit reports. Four Golden Signals /metrics endpoint.
Governed Agent Stack
FloorMind turns a plain-English factory question into SQL that sql-sop lints, query-warden gates by role, pii-veil masks, and agent-blackbox chains into a tamper-evident ledger, while elenchus runs governed surveys alongside. The whole stack is assembled into one monorepo and merged into a single Next.js console: chat, factory KPIs, compliance, waste, documents, surveys, audit and configuration, with every backend proxied server-side so the browser never touches CORS or auth headers.
- sql-explorer-mcp — read-only MCP server for SQL Server, Postgres and SQLite behind a three-layer safety stack (PyPI)
- pr-sop — PR governance checks (changelog drift, version consistency) as a CLI or GitHub Action (PyPI)
- morning-brief — rule-based Gmail triage, zero LLM, read-only OAuth (PyPI)
- MediAsk — hackathon health Q&A app: NHS-verified guidance, 18 languages, voice input
My Stack
Tools I reach for daily, grouped by what they're for.
Languages
Data Engineering
AI & Agents
Web & API
BI & Visualisation
Infra & DevOps
Manufacturing & Compliance Domain
Open Source Contributions
Maintainer or substantive contributor — not just docs typo fixes.
drt-hub/drt
Collaborator on the multi-source data sync engine across three releases. v0.5: destination connectors including Mixpanel and the official connector tutorial. v0.6: --threads N parallel orchestration with a thread-safe StateManager and 11 parallel-dispatch tests. v0.7: --quiet for CI/cron use cases. Plus reviewer voice on the json_columns config PR where the early-validation suggestion shaped the final implementation.
sql-sop
Three outside contributors have merged nine lint rules into it. Six came from @mvanhorn (W019, W016, W015, W023, W021, W012), with W011 and P005 from @tmchow and the OVER() window check from @Prabhu-1409. I run the review queue, publish to PyPI via Trusted Publishing (2,000+ installs, mirrors excluded), maintain the security and governance policy, and keep a public ROADMAP plus a one-file scaffold so a new contributor can add a rule without touching the core.
pr-sop
Shipped v0.1.0, v0.1.1 (third-party rev: pin false-positive fix), and v0.1.2 (CI-merge-commit tag lookup fix) to PyPI in 24 hours. Full governance, security, contributing, and code-of-conduct documents published.
Signature Upstream Pull Requests
Merged fixes and features in the SQL, data and AI tools I use every day.
tobymao/sqlglot#7824
Presto/Trino mapped native SHA256/512 to the string-hash expression, silently changing hash values in transpilation. Traced the root cause into the transpiler that underpins dbt and SQLMesh, mirrored the existing MD5 handling, and it merged the next morning.
apache/superset#39118
Aligned the dashboard download permission with the explore path. Refined to a single-responsibility change after maintainer review, then merged into Apache Superset.
sqlfluff/sqlfluff#8088
A reported linter false-positive on quoted T-SQL aliases turned out to be a dialect-level bug: the patched single_quote lexer had dropped normalization for every single-quoted token. Fixed at the root with a regression test; all 545 dialect fixtures pass unregenerated.
drt-hub/drt#668 (+ #608, #678)
Collaborator across releases. Shipped sync.mask (hash / redact / truncate PII before load) as a pure transform at the field-mapping seam, plus the Mixpanel destination connector and the official connector tutorial.
tobymao/sqlglot#7832
Follow-up to #7824, opened at the maintainer's invitation: mapped BigQuery's native SHA512 to the digest expression and gave Presto/Trino a type-aware encode, keeping byte-for-byte hash semantics across six dialects.
openfoodfacts/openfoodfacts-server#13892
A one-line fix to a safety-critical open dataset used worldwide: "groundnut", the standard UK and Indian term, was missing from the peanut allergen synonyms, so allergen cross-checks could silently miss it.
- SQLMesh · #5888 — hex-string surrogate keys for SHA256/512 on Presto/Trino
- sops · #2262 — stop requiring a matching creation rule when a key is given explicitly
- dlt · #4026 — resolve GCP service-account against oauth credentials correctly
- Octopus Energy · tentaclio#243 — expand
~in the credentials file path - textual-fastdatatable · #154 — escape and truncate bytes values in the cell formatter
- drt · #902 —
drt servedelivery contract: coalescing, 202 + run id, pluggable auth
Activity
What the contribution graph and language mix look like right now.
Get in Touch
Open to data engineering, data ops, manufacturing analytics, and on-prem AI roles. Especially interested in roles where the data domain genuinely matters and the system has to keep running on a Sunday night.
// COMMUNITY
discord.gg/gBr77yYPkDHow I work
Three honest stages — same flow whether the work is a pipeline, a dashboard, or shipping a tool to PyPI.
Discovery
Read the source. Read the data. Talk to the people who actually run the line. Confirm the problem is the problem before writing code.
Architecture
Sketch the smallest thing that delivers value. Pick boring, well-documented tools. Write the tests first when the contract matters.
Delivery
Ship in CI behind a green build. Document the trade-offs. Set up a way for the next person (sometimes me) to understand it without paging me.