feat(nyc-taxi): add ML entities to the taxi sample dataset - #223
Open
TommyTranX wants to merge 1 commit into
Open
TommyTranX wants to merge 1 commit into
TommyTranX wants to merge 1 commit into
Conversation
None of the sample datasets in this repo ship ML entities, so there is nowhere
to exercise DataHub's ML metadata model against realistic data - anyone demoing
or testing mlModel / mlFeature / mlFeatureTable has to invent fixtures first.
add_ml_entities.py builds the ML half of the existing pipeline:
staging_trips --(DerivedFrom)--> mlFeature x6 --(Consumes)--> mlModel
|
v
mlModelDeployment
Six rolling 7-day features, each sourced from a column that actually exists in
staging_trips, a feature table, a model, and a deployment. Entities are
namespaced by platform instance so nyc_taxi and nyc_taxi_pipeline coexist.
Run after add_lineage.py; the model's upstream chain then reaches raw_trips:
searchAcrossLineage(mlModel, UPSTREAM)
degree 1 mlFeature x6
degree 2 staging_trips
degree 3 raw_trips
Paired with nyc_taxi_pipeline.db this makes the planted staleness reachable
from a model - a stale table upstream of a production model's features, which
is the usual shape of a silent ML failure and is not otherwise testable with
the shipped fixtures.
Two implementation notes:
- mlFeature.sources accepts dataset URNs only (relationship annotation is
entityTypes:["dataset"]); a schemaField URN is rejected with "is not a valid
destination". The originating column is kept as a source_column custom
property.
- URNs are constructed rather than discovered via search. add_lineage.py and
add_metadata.py use the search API, which is populated asynchronously, so
they can print "No datasets found" immediately after a successful ingest.
Building URNs from the platform instance and table name avoids depending on
index freshness.
Follows the existing scripts' conventions: --instance=, --all, --dry-run,
--help, and the same output style. Verified end to end against DataHub OSS
quickstart v1.5.0.6 with both variants.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
None of the sample datasets in this repo ship ML entities. There is no
mlModel,mlFeatureormlFeatureTableanywhere indatasets/, so anyone demoing or testingDataHub's ML metadata model has to invent their own fixtures before they can start.
This adds them on top of the taxi pipeline that already exists.
What
add_ml_entities.pybuilds the ML half of the pipeline:mlFeatureTable<instance>_taxi_featuresmlFeature× 6trips_7d,avg_fare_7d,avg_distance_7d,avg_duration_7d,passenger_mean_7d,revenue_7dmlModel<instance>_taxi_demand_forecastmlModelDeployment<instance>_taxi-demand-prodEvery feature is sourced from a column that actually exists in
staging_trips. Entitiesare namespaced by platform instance, so
nyc_taxiandnyc_taxi_pipelinecoexist.Why it is worth having alongside
nyc_taxi_pipeline.dbRun after
add_lineage.py, the model's upstream chain reaches the raw table:Paired with the staleness variant, that makes the planted defect reachable from a
model — a stale table sitting upstream of a production model's features. That is the
usual shape of a silent ML failure, and it is not testable with the fixtures as they
stand today.
Implementation notes
mlFeature.sourcesaccepts dataset URNs only. Its relationship annotation isentityTypes: ["dataset"], and GMS rejects aschemaFieldURN with "is not a validdestination". The originating column is kept as a
source_columncustom propertyinstead. Worth knowing for anyone who assumes column-level sourcing is available here.
URNs are constructed, not discovered by search.
add_lineage.pyandadd_metadata.pyresolve URNs through the search API. DataHub's search index ispopulated asynchronously by the MAE consumer, so those scripts can print
No datasets found (run ingestion first)immediately after a successful ingest — I hitthis while testing. Building URNs from the platform instance and table name removes the
dependency on index freshness. I did not change the existing scripts in this PR, but the
same fix would apply to them if that is wanted.
Conventions
Follows the existing scripts:
--instance=,--all,--dry-run,--help, the same✓ / ✗ / ⚠output style, and the sameDATAHUB_SERVERconstant. README updated with asection documenting the entities and the run order.
Testing
Verified end to end against DataHub OSS quickstart v1.5.0.6:
--dry-runand--all --dry-run— no writes--instance=nyc_taxi_pipelineafterdatahub ingest -c ingest_pipeline.yamlandadd_lineage.py— all 9 aspects emitted, entities render in the UIDOWNSTREAMfromstaging_tripsreaches the sixfeatures at degree 1 and the model at degree 2;
UPSTREAMfrom the model reachesraw_tripsat degree 3Found and built while working on the DataHub Agent Hackathon.
Related: #222 (the README's documented defect values don't match the shipped database).