This repository handles the HTTP Archive data pipeline, which takes the results of the monthly HTTP Archive run and saves this to the httparchive dataset in BigQuery.
Pipelines are transformed via Dataform and orchestrated by Google Cloud Composer (Apache Airflow) in GCP. Airflow DAGs reside in airflow/dags/ and are synced to Cloud Composer on merge to main.
| Pipeline | DAG | Dataform Tags | Primary Outputs |
|---|---|---|---|
| HTTP Archive Crawl | airflow/dags/crawl_complete.py |
crawl_complete, crawl_complete_reports |
httparchive.crawl.* (Analytics Hub), httparchive.blink_features.usage (chromestatus.com) |
| Technology Report | airflow/dags/crux_ready.py |
crux_ready, crux_ready_reports |
httparchive.reports.cwv_tech_*, httparchive.reports.tech_* (Tech Report) |
For complete orchestration details (triggers, sensors, pre-tasks) and development workspace guidelines, see Dataform Documentation. For overall GCP infrastructure and data flows, see Infrastructure Overview.
-
Install dependencies:
npm install
-
Available Scripts:
npm run format- Format code and fix Markdown issuesnpm run lint- Run linting checks on Markdown files, and compile Dataform configs
This repository uses:
- Markdownlint for Markdown file formatting
- Dataform's built-in compiler for SQL validation