Skip to content

feat(nyc-taxi): add ML entities to the taxi sample dataset - #223

Open
TommyTranX wants to merge 1 commit into
datahub-project:mainfrom
TommyTranX:add-ml-entities-nyc-taxi
Open

TommyTranX wants to merge 1 commit into
datahub-project:mainfrom
TommyTranX:add-ml-entities-nyc-taxi

Conversation

@TommyTranX

Copy link
Copy Markdown

Why

None of the sample datasets in this repo ship ML entities. There is no mlModel,
mlFeature or mlFeatureTable anywhere in datasets/, so anyone demoing or testing
DataHub's ML metadata model has to invent their own fixtures before they can start.

This adds them on top of the taxi pipeline that already exists.

What

add_ml_entities.py builds the ML half of the pipeline:

staging_trips ──(DerivedFrom)──▶ mlFeature × 6 ──(Consumes)──▶ mlModel ──▶ mlModelDeployment
Entity Name
mlFeatureTable <instance>_taxi_features
mlFeature × 6 trips_7d, avg_fare_7d, avg_distance_7d, avg_duration_7d, passenger_mean_7d, revenue_7d
mlModel <instance>_taxi_demand_forecast
mlModelDeployment <instance>_taxi-demand-prod

Every feature is sourced from a column that actually exists in staging_trips. Entities
are namespaced by platform instance, so nyc_taxi and nyc_taxi_pipeline coexist.

Why it is worth having alongside nyc_taxi_pipeline.db

Run after add_lineage.py, the model's upstream chain reaches the raw table:

searchAcrossLineage(mlModel, UPSTREAM)
  degree 1  mlFeature × 6
  degree 2  staging_trips
  degree 3  raw_trips

Paired with the staleness variant, that makes the planted defect reachable from a
model
— a stale table sitting upstream of a production model's features. That is the
usual shape of a silent ML failure, and it is not testable with the fixtures as they
stand today.

Implementation notes

mlFeature.sources accepts dataset URNs only. Its relationship annotation is
entityTypes: ["dataset"], and GMS rejects a schemaField URN with "is not a valid
destination"
. The originating column is kept as a source_column custom property
instead. Worth knowing for anyone who assumes column-level sourcing is available here.

URNs are constructed, not discovered by search. add_lineage.py and
add_metadata.py resolve URNs through the search API. DataHub's search index is
populated asynchronously by the MAE consumer, so those scripts can print
No datasets found (run ingestion first) immediately after a successful ingest — I hit
this while testing. Building URNs from the platform instance and table name removes the
dependency on index freshness. I did not change the existing scripts in this PR, but the
same fix would apply to them if that is wanted.

Conventions

Follows the existing scripts: --instance=, --all, --dry-run, --help, the same
✓ / ✗ / ⚠ output style, and the same DATAHUB_SERVER constant. README updated with a
section documenting the entities and the run order.

Testing

Verified end to end against DataHub OSS quickstart v1.5.0.6:

  • --dry-run and --all --dry-run — no writes
  • --instance=nyc_taxi_pipeline after datahub ingest -c ingest_pipeline.yaml and
    add_lineage.py — all 9 aspects emitted, entities render in the UI
  • Lineage verified in both directions: DOWNSTREAM from staging_trips reaches the six
    features at degree 1 and the model at degree 2; UPSTREAM from the model reaches
    raw_trips at degree 3

Found and built while working on the DataHub Agent Hackathon.
Related: #222 (the README's documented defect values don't match the shipped database).

None of the sample datasets in this repo ship ML entities, so there is nowhere
to exercise DataHub's ML metadata model against realistic data - anyone demoing
or testing mlModel / mlFeature / mlFeatureTable has to invent fixtures first.

add_ml_entities.py builds the ML half of the existing pipeline:

  staging_trips --(DerivedFrom)--> mlFeature x6 --(Consumes)--> mlModel
                                                                  |
                                                                  v
                                                        mlModelDeployment

Six rolling 7-day features, each sourced from a column that actually exists in
staging_trips, a feature table, a model, and a deployment. Entities are
namespaced by platform instance so nyc_taxi and nyc_taxi_pipeline coexist.

Run after add_lineage.py; the model's upstream chain then reaches raw_trips:

  searchAcrossLineage(mlModel, UPSTREAM)
    degree 1  mlFeature x6
    degree 2  staging_trips
    degree 3  raw_trips

Paired with nyc_taxi_pipeline.db this makes the planted staleness reachable
from a model - a stale table upstream of a production model's features, which
is the usual shape of a silent ML failure and is not otherwise testable with
the shipped fixtures.

Two implementation notes:

- mlFeature.sources accepts dataset URNs only (relationship annotation is
  entityTypes:["dataset"]); a schemaField URN is rejected with "is not a valid
  destination". The originating column is kept as a source_column custom
  property.

- URNs are constructed rather than discovered via search. add_lineage.py and
  add_metadata.py use the search API, which is populated asynchronously, so
  they can print "No datasets found" immediately after a successful ingest.
  Building URNs from the platform instance and table name avoids depending on
  index freshness.

Follows the existing scripts' conventions: --instance=, --all, --dry-run,
--help, and the same output style. Verified end to end against DataHub OSS
quickstart v1.5.0.6 with both variants.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant