Skip to content

Repository files navigation

HTTP Archive datasets pipeline

This repository handles the HTTP Archive data pipeline, which takes the results of the monthly HTTP Archive run and saves this to the httparchive dataset in BigQuery.

Pipelines & Orchestration

Pipelines are transformed via Dataform and orchestrated by Google Cloud Composer (Apache Airflow) in GCP. Airflow DAGs reside in airflow/dags/ and are synced to Cloud Composer on merge to main.

Pipeline DAG Dataform Tags Primary Outputs
HTTP Archive Crawl airflow/dags/crawl_complete.py crawl_complete, crawl_complete_reports httparchive.crawl.* (Analytics Hub), httparchive.blink_features.usage (chromestatus.com)
Technology Report airflow/dags/crux_ready.py crux_ready, crux_ready_reports httparchive.reports.cwv_tech_*, httparchive.reports.tech_* (Tech Report)

For complete orchestration details (triggers, sensors, pre-tasks) and development workspace guidelines, see Dataform Documentation. For overall GCP infrastructure and data flows, see Infrastructure Overview.

Development Setup

  1. Install dependencies:

    npm install
  2. Available Scripts:

    • npm run format - Format code and fix Markdown issues
    • npm run lint - Run linting checks on Markdown files, and compile Dataform configs

Code Quality

This repository uses:

  • Markdownlint for Markdown file formatting
  • Dataform's built-in compiler for SQL validation

About

The data pipeline for HTTP Archive orchestrated by Dataform

Resources

Security policy

Stars

7 stars

Watchers

4 watching

Forks

Used by

Contributors

Languages