Mozilla Data Documentation
1.
Introduction
1.1.
What Data does Mozilla Collect?
1.2.
Tools for Data Analysis
1.3.
Terminology
2.
Tutorials & Cookbooks
2.1.
Getting Started
2.1.1.
Gaining Access
2.1.2.
Getting Help
2.1.3.
Reporting a problem
2.2.
Analysis
2.2.1.
Data Discovery Tools
2.2.1.1.
Using the Data Catalog
2.2.1.2.
Using the Glean Dictionary
2.2.1.3.
Using the Probe Dictionary
2.2.2.
Data Monitoring - Intro to Bigeye
2.2.2.1.
Interface
2.2.2.2.
Deploying Metrics
2.2.2.3.
Collections
2.2.2.4.
Issues Management
2.2.2.5.
Cost Considerations
2.2.2.6.
bigquery-etl and Bigeye
2.2.2.7.
Further Reading
2.2.3.
Data Modeling
2.2.3.1.
Where to store the data model assets
2.2.3.2.
Using aggregates in BigQuery and Looker
2.2.3.3.
Shredder mitigation
2.2.4.
Working with Looker
2.2.4.1.
Introduction to Looker
2.2.4.2.
Normalizing Country Data
2.2.4.3.
Normalizing Browser Version Data
2.2.4.4.
Using Growth and Usage Dashboards
2.2.4.5.
Using the Event Counts Explore
2.2.4.6.
Using the Funnel Analysis Explore
2.2.4.7.
Looker Performance - Caching
2.2.5.
Other Data Analysis Tools
2.2.5.1.
Introduction to GLAM
2.2.5.2.
Introduction to Operational Monitoring
2.2.5.3.
Introduction to STMO
2.2.6.
Accessing Public Data
2.2.7.
Accessing and working with BigQuery
2.2.7.1.
Access
2.2.7.2.
Writing Queries
2.2.7.3.
Optimization
2.2.7.4.
Accessing Desktop Data
2.2.7.5.
Accessing Glean Data
2.2.7.6.
Accessing Additional Properties
2.2.7.7.
Custom analysis with Spark
2.2.8.
Dataset-Specific
2.2.8.1.
Working with Normandy events
2.2.8.2.
Working with Crash Pings
2.2.8.3.
Working with Bit Patterns in Clients Last Seen
2.2.8.4.
Visualizing Percentiles of a Main Ping Exponential Histogram
2.2.9.
Real-time
2.2.9.1.
Working with Live Data
2.2.9.2.
Seeing Your Own Pings
2.2.9.3.
See Real-time search metrics
2.2.10.
Metrics
2.3.
Operational
2.3.1.
Creating a Prototype Data Project on Google Cloud Platform
2.3.2.
Creating Static Dashboards with Protosaur
2.3.3.
Scheduling Queries
2.3.4.
Building and Deploying Containers to GCR with CircleCI
2.3.5.
Publishing Datasets
2.3.6.
Connecting Sheets and External Data to BigQuery
2.4.
Sending telemetry
2.4.1.
Implementing Experiments
2.4.2.
Sending Events
2.4.3.
Sending a Custom Ping
3.
Data Platform Reference
3.1.
Data Stack Overview
3.2.
Guiding Principles for Data Infrastructure
3.3.
Glean overview
3.4.
Overview of Mozilla's Data Pipeline
3.4.1.
HTTP Edge Server Specification
3.4.2.
Event Pipeline Detail
3.4.3.
Schemas
3.4.4.
Glean Data
3.4.5.
Channel Normalization
3.4.6.
Sampling
3.4.7.
Filtering
3.4.8.
BigQuery Artifact Deployment
3.5.
Common Analysis Gotchas
3.6.
SQL Style Guide
3.7.
Airflow Gotcha's
3.8.
Telemetry Behavior Reference
3.8.1.
History of Telemetry
3.8.2.
Profile Behavior
3.8.2.1.
Profile Creation
3.8.2.2.
Real World Usage
3.8.2.3.
Profile History
3.8.3.
Engagement metrics
3.8.4.
User states/Segments
3.9.
Experimentation
3.10.
Metric Hub
3.11.
External data integration using Fivetran
3.12.
Project Glossary
4.
Dataset Reference
4.1.
Pings
4.2.
Derived Datasets
4.2.1.
Active Profiles
4.2.2.
Active Users
4.2.3.
Addons
4.2.4.
Addons Daily
4.2.5.
Autonomous System Aggregates
4.2.6.
Clients Daily
4.2.7.
Clients Last Seen
4.2.8.
Events
4.2.9.
Events Daily
4.2.10.
Firefox Android Clients
4.2.11.
Main Ping Tables
4.2.12.
Main Summary