Repository navigation
Adapt to risk vax burden protocol - #2
Conversation
…ort_id (target_cohort) / Disease-specific variables by target
…ot, and reorganize using protocol headings
|
Hi Martina, great to see this moving. Some quick thoughts from me before I dig into the code: Added a cohort_id parameter Nice idea. Consolidated everything into a single dataset_definition I think this is sensible and pragmatic. It might make the whole pipeline overall a bit slower, and total size of the outputted datasets larger, but it simplifies the code and is easier to work with. However:
Unfortunately, recording dates for ethnicity are very unreliable due to patient registration movement through practices - the date of recording is lost during any transfer of records (see https://doi.org/10.1186/s12916-024-03499-5) and it defaults to 1900-01-01 (or something). So even if you choose to ignore future-dated ethnicity codes, you still might be picking up ethnicity codes that were recorded in the future (just set to 1900 by default following a move). This introduces a different type of bias (more likely to have an ethnicity code for patients who have moved in the future). Extracting ethnicity just once for each patient will also significantly decrease runtime. So something to reverting back to if speed is an issue. Simplified vaccination history Again, sensible and pragmatic. But it does make it difficult to identify and deal with any vaccine data quality issues of the sort we've seen in the OVERTURE work. Ideally, we would be able to create a single "cleaned_vaccinations" table view that could be queried like any other table, but we're not there yet. We can consider later whether data quality issues are concerning enough to use ELD to deal with directly. Age calculated the day before cohort start I would choose the day of the cohort start, because date of birth is rounded to the first of the month. If we choose the first of the month as the cohort start date, and the day before that as the age-at date, then we're essentially choosing the first of the prior month as the age-at date (i think??). Also, we chose the age-at date months after the cohort start date in ECHO, because that was how age-based vaccine eligibility worked (if you were old enough at any point during the campaign, then you were considered eligible). I don't know whether that's a sensible choice here too, but will defer to Ed. |
Hi @eparker12 and @wjchulme,
I created a first draft adapting the vaccine history scripts for the Harmonised Assessment of Risk Groups for Vaccine Prioritisation analysis.
Main changes
project.yamlcohort_idparameter to identify each cohort (one per target disease and time period; e.g.flu_2023_24corresponds to the 2023/24 influenza campaign). I think this might be better than 2 parameters (e.g. target_disease + time_period), so we can run each season independently. But I am open to this approach if it is more efficient.design.Rcohort_end_date.campaign_infowithcohort_info.dataset_definition.pyConsolidated everything into a single
dataset_definition.Simplified vaccination history by removing event-level vaccination data and instead deriving two dates:
Removed the separate fixed dataset. We discussed with Ed that using only the latest recorded ethnicity could introduce bias, as ethnicity recording has improved over time and individuals who survive longer have more opportunities to have their ethnicity recorded. After removing ethnicity, very few truly fixed variables remained, so I incorporated them directly into the main dataset definition.
Added disease-specific vaccination and outcome variables, using an
ifstatement indataset_definition.pyto generate the appropriate variables based on thecohort_idparameter.I still need to add the mild outcomes. If we consider this is OK I can move to do that.
Age calculated the day before cohort start (
cohort_start_date - days(1))prepare.Rbaseline_vax_statusvariable following the protocol definitions for each vaccine.report_cohort.Rreport_cohort.Rreport_snapshot, and reorganised it to follow the protocol structure.There are still a few things to tidy up, but I think the overall structure is now much closer to the protocol and should make future extensions easier.