A general database builder on top of pypath.
make setupThis project uses uv for dependency management.
The omnipath_build pipeline streams inputs_v2 sources into DuckDB evidence
tables, canonicalizes them in DuckDB, and copies the projected evidence and
canonical rows into PostgreSQL.
For the data model, phase boundaries, refresh semantics, and common workflows, see docs/pipeline.md.
For tuning the build's Postgres memory and parallelism (and how to size it for your deployment, including the lab's docker memory cap), see docs/build-tuning.md.
The default database URL is:
postgresql://omnipath:omnipath@localhost:55432/omnipathOverride it with DATABASE_URL=... when needed.
Build local resolver parquet files:
make resolverLimit resolver builds for smoke tests:
make resolver MAX_RECORDS=100000Use a single PubChem SDF shard during development:
make resolver PUBCHEM_URL=https://example.org/pubchem.sdf.gzCreate schema and supporting indexes:
make db-setupStart from a clean schema. This defers secondary evidence indexes until canonicalization, so ingest is faster:
make db-setup DROP_EXISTING=1Use another schema:
make db-setup SCHEMA=omnipath_test DROP_EXISTING=1Drop and recreate the target schema without loading resolver tables:
make db-resetLoad any sources that are not already present. The DuckDB/PostgreSQL loader projects evidence, canonicalizes it, and copies the result into PostgreSQL:
make loadLoad one source if it is not already present:
make load SOURCE=bindingdbLoad multiple missing sources:
make load SOURCES=uniprot,bindingdb,intactUse a shared staging worker pool:
make load LOAD_JOBS=5LOAD_JOBS is used by the staged loader as one shared pool. Workers are
assigned across sources first; when preparse shards are available, idle workers
can stage shards from the same source while PostgreSQL COPY continues to run
serially.
Large staged loads keep simple preparse parquet shards under
pypath-data/<source>/preparse/. These shards are reused by default and are
rebuilt when FORCE_REFRESH=1 is set.
Refresh existing source content by deleting it first, then loading current parser output:
make reload SOURCE=bindingdbRun make derive after loading the selected sources to refresh query indexes,
derived count/search tables, and bitmaps.
Build resolver files, recreate the database schema, load all sources through the DuckDB/PostgreSQL pipeline, and derive query tables:
make all DROP_EXISTING=1If resolver files already exist, run:
make db-setup DROP_EXISTING=1
make load
make deriveLimit source rows per dataset:
make load SOURCE=bindingdb MAX_RECORDS=200000 SCHEMA=omnipath_testThe default load batch size is BATCH_SIZE=50000, which is safer
for full loads.
Reset omnipath_build content tables without dropping resolver tables:
make reset-contentRun the Python tests:
uv run pytest tests -qCheck table sizes when PostgreSQL runs in Docker:
docker exec -i omnipathv2-main-mmvxvb-omnipathv2-postgres-1 \
psql -U omnipath -d omnipath -v ON_ERROR_STOP=1 <<'SQL'
SELECT
table_name,
pg_size_pretty(pg_total_relation_size(format('%I.%I', table_schema, table_name)::regclass)) AS total_size,
pg_size_pretty(pg_relation_size(format('%I.%I', table_schema, table_name)::regclass)) AS table_size,
pg_size_pretty(pg_indexes_size(format('%I.%I', table_schema, table_name)::regclass)) AS indexes_size
FROM information_schema.tables
WHERE table_schema = 'public'
AND table_type = 'BASE TABLE'
ORDER BY pg_total_relation_size(format('%I.%I', table_schema, table_name)::regclass) DESC;
SQL