Highways & Railtracks
AWS Registry of Open Data
Datadory delivers aws registry of open data data covering 1,069 catalogued dataset entries curated by Amazon Web Services - the OpenStreetMap planet archive with minutely change replication and weekly snapshots, cloud-native geospatial conversions in Parquet and Zarr, and partner collections from NOAA, NASA and Digital Earth Africa - normalised into one documented feed for highway and rail-track work.
API, files, or your warehouse. Daily, weekly, or hourly.
- Where it covers
- Global - entries span worldwide and regional datasets across many domains, including global OSM archives relevant to highways and rail
- How far back
- Registry continuously extended with a public changelog of recent additions; member datasets range from static archives to minutely replication streams
- How fine
- One registry entry per dataset, with per-resource detail blocks for each bucket, topic or API behind it
What is the AWS Registry of Open Data?
AWS Registry of Open Data is the highways & railtracks data hub catalog's widest record - not one dataset but a curated directory of 1,069 of them, observed on the registry's listing during the August 2026 research pass. Amazon Web Services curates the shelf; third parties stock it. Every entry follows the same template: a title and tag line, a short description, usage examples with citations, tutorials (many paired with SageMaker Studio Lab notebooks), and a Resources-on-AWS block spelling out each underlying resource - type, bucket ARN, region, and a ready-to-run command.
The partner collections give the catalog its breadth: the EPA, Allen Institute for AI, Biohub, Digital Earth Africa, Data for Good at Meta, NIH STRIDES, NOAA's Open Data Dissemination Program, the Space Telescope Science Institute under a Space Act Agreement, and the Amazon Sustainability Data Initiative. For highway and rail-track work, the anchor entry is OpenStreetMap on AWS - OSMF-managed planet and history PBFs, changeset and discussion XML, and OsmChange replication files updated at minute/hour/day cadence with weekly snapshots, replicated across eu-central-1 and us-west-2. A 'daylight-osm' entry sits alongside it in the same domain.
One caveat travels with every row: the registry states that datasets are provided and maintained by a variety of third parties under a variety of licenses - AWS hosts almost none of the content itself. Datadory delivers the transport-relevant entries as shaped extracts with per-entry terms confirmed at sample request.
What do sample rows look like?
Two records off the transport-relevant core - first the registry-level metadata entry for OpenStreetMap on AWS, then two of its per-resource blocks:
record_type : dataset_entry
title : OpenStreetMap on AWS
tags : disaster response; geospatial; mapping; osm
update_freq : minutely/hourly/daily (changes), weekly (snapshots)
managed_by : OpenStreetMap Foundation (OSMF) and Pacific Atlas
record_type : resource_block
resource_type: S3 Bucket
arn_region : arn:aws:s3:::osm-planet-eu-central-1 | eu-central-1
description : Primary - planet & history PBFs, changeset/discussion XML,
OsmChange replication files
record_type : resource_block
resource_type: S3 Bucket
arn_region : arn:aws:s3:::osm-planet-us-west-2 | us-west-2
description : Replicated copy of planet & history PBFs + replication filesRead the anatomy rather than the buckets. The dataset-entry record carries the who-and-how-often layer: named maintainers, tag taxonomy, stated refresh cadence. Each resource-block record carries the where-and-what layer: an ARN pinning one physical store to one region, plus a plain-language inventory of its contents. That two-level shape repeats across all 1,069 entries, which is what makes programmatic sweeps of the whole catalog practical.
What fields does the dataset include?
Eight verified fields describe every registry-level entry, captured verbatim during the August 2026 research pass:
title- the dataset name as listed; 'OpenStreetMap on AWS' is the transport anchor.Description- short statement of what the member dataset contains.Update Frequency- the maintainer's own stated cadence, from static archives to minutely replication streams.License- pointer to the entry's own license terms, set by its provider rather than by AWS.Documentation- upstream docs or the open-data-docs tree maintained alongside the registry.Managed By- the organization answerable for the content, e.g. OpenStreetMap Foundation (OSMF) and Pacific Atlas.Resource type / ARN / Region- the per-resource detail block behind the entry: S3 bucket, SNS topic or API, its ARN, its region.AWS CLI Access- the ready-to-run command published with the entry, typically flagged for anonymous reads.
The resource-block columns (type, ARN, region, description), per-entry usage-example citations and tutorial links fold under additional fields on request.
Which extras arrive only on request?
The eight-field core covers the entry level; the deeper layers ship with your sample rather than padding this page:
- Per-resource block columns. Resource type, Amazon Resource Name, AWS region and the resource's own description for every bucket, topic and API behind an entry.
- Usage examples and tutorials. The citations and notebook links attached to each entry, including the SageMaker Studio Lab pairings common across the catalog.
- Collection-level metadata. The ten-plus partner collections - from NOAA's dissemination program to Digital Earth Africa - documented so you can sweep a collection instead of entry by entry.
What does coverage look like across geography, time and granularity?
Geography - global, twice over. The catalog spans worldwide and regional datasets across many domains, and its transport anchor is itself planetary: OSM planet archives cover every road and rail way mapped worldwide, replicated across two regions for resilience.
Temporal - the directory keeps growing, with a public changelog tracking recent additions. Member datasets run the full range underneath: static historical archives at one end, minutely change replication on the other, weekly snapshots in between. That spread means a corridor study and a live-monitoring build can draw on the same catalog without either waiting on the other.
Granularity - one registry entry per dataset, with per-resource detail blocks for each bucket, topic or API behind it. The grain is deliberate: metadata about datasets, uniform enough to parse, pointing at resources large enough to matter - individual member datasets range from megabytes to petabytes.
How is the data delivered?
API, files, or your warehouse. Daily, weekly, or hourly.
Name the entries or collections you need - the OSM planet spine, a NOAA archive, a whole partner collection - and the sample arrives shaped to that scope with the field dictionary attached. Because the catalog is a directory, delivery is a selection problem before it is a synchronization one: we track which entries exist and how they are described, you consume the result as files, a feed or warehouse tables on the cadence your workflow asks for. Per-entry terms travel with the extract.
Who uses this data, and for what?
- Road-network extraction at global scale - planet PBFs plus minutely change files give a base map that never ages out: extract road geometry once, then apply deltas; see data scientists use cases
- Asset-location enrichment - cross bridges, depots or site lists against OSM way geometry to attach road class, surface and access attributes without a survey
- Change-detection studies - OsmChange replication supports before/after comparison of network edits around construction corridors or disaster zones; see developers & builders use cases
- Cloud-warehouse geospatial loading - Parquet, COG and Zarr conversions land in columnar engines without format-conversion plumbing
- Portfolio scanning across domains - one entry template across 1,069 listings lets a single parser sweep climate, imagery and mobility archives together
- Infrastructure scoping and due diligence - consultants frame corridor questions against measured network context rather than press releases; see market researchers use cases
Which personas get the most value?
Data scientists get a planet-scale training corpus for road-segmentation and routing models, with change feeds that support continual-learning experiments rather than one-shot training runs. Developers and builders get stable identifiers and a uniform entry template - sweeping the full catalog programmatically is a day's work, not a quarter's. Journalists, academics and students get named maintainers and citable usage examples per entry, which keeps provenance clean in published work. Market researchers and consultants scoping infrastructure-adjacent questions get one catalog answering across climate, transport and imagery domains instead of ten separate hunts.
How does it compare to alternatives in its slice?
Within highways & railtracks data, this record owns breadth through aggregation: 1,069 entries under one template, spanning far more domains than any single-domain rival. Its neighbours own depth instead. Overture Maps Transportation Theme re-models road and rail geometry as an entity graph with consistent IDs - cleaner joins, narrower scope. The OpenStreetMap Overpass API queries OSM selectively when hauling the planet archive would be waste. Geofabrik extracts cut OSM to country or region size for corridor-scale work. USDOT's portal stays authoritative for US DOT series, and the National Bridge Inventory measures structures this catalog merely points at. Pair them: the registry for reach, the specialists for measurement.
What should I know before requesting a sample?
Notes worth having in hand:
- A directory, not a dataset - the record catalogs 1,069 member datasets; quality and relevance depend entirely on which entries serve your question.
- Third-party maintenance is the norm - each entry names its own maintainer and carries its own license terms; per-entry terms are confirmed with your sample.
- Search lags the listing - during research the inline search reported 13 matches while the page linked 1,069 detail pages, so enumeration beats search for completeness.
- The OSM entry is the transport spine - minutely change replication, weekly snapshots, two-region redundancy, cloud-native conversions alongside the PBFs.
Where breadth stops, neighbours pick up: Overture for entity-modelled roads, Overpass for selective queries, Geofabrik for regional extracts, USDOT and NBI for measured US series. Name the entries, collections or domains you need and the sample returns shaped to them with the complete field dictionary attached.
Field dictionary
Every field below is documented against real records. The full dictionary ships with the sample.
| field | type | definition | example |
|---|---|---|---|
title | string | Dataset name as listed in the registry. | OpenStreetMap on AWS |
Description | text | Short description of what the member dataset contains. | Planet PBFs and change files for the global OSM map |
Update Frequency | string | Stated refresh cadence of the underlying data, set by its maintainer. | minutely/hourly/daily (changes), weekly (snapshots) |
License | string | Pointer to the entry's own license terms, set by its third-party provider. | OSM copyright notice |
Documentation | string | Link to upstream docs or the open-data-docs tree maintained alongside the registry. | open-data-docs OSM PDS page |
Managed By | string | Organization responsible for the dataset's content and upkeep. | OpenStreetMap Foundation (OSMF) and Pacific Atlas |
Resource type / ARN / Region | text | Per-resource block: AWS resource type (S3 bucket, SNS topic), its Amazon Resource Name and region. | arn:aws:s3:::osm-planet-eu-central-1 | eu-central-1 |
AWS CLI Access | string | Ready-to-run command published with each entry, flagged by its maintainer for anonymous or credentialed reads. | One per entry, in its Resources block |
What teams do with it
- Road-network extraction at global scale Planet PBFs plus minutely change files give a base map that never ages out - extract road geometry once, then apply only the deltas.
- Asset-location enrichment Cross bridges, depots or site lists against OSM way geometry to attach road class, surface and access attributes without a survey.
- Change-detection studies OsmChange replication streams support before/after comparisons of network edits around construction corridors or disaster zones.
- Cloud-warehouse geospatial loading Parquet, COG and Zarr formats land in columnar engines directly, skipping format-conversion plumbing.
- Portfolio scanning across domains One entry template across all 1,069 listings lets a single parser sweep climate, imagery and mobility archives.
Questions buyers ask
How many datasets are in the AWS Registry of Open Data?
The listing carried 1,069 distinct dataset detail links when observed during the August 2026 research pass. Entries span geospatial, climate, biology, imagery and mobility domains, grouped into curated partner collections. The transport anchor is OpenStreetMap on AWS with planet archives and minutely replication; Datadory delivers whichever entries serve your question as shaped extracts.
What is the OpenStreetMap on AWS entry?
The transport-relevant anchor of the catalog: OSMF-managed planet and history PBFs in protobuf format, changeset and discussion XML, and OsmChange replication files updated at minute/hour/day cadence with weekly snapshots. Content replicates across eu-central-1 and us-west-2 buckets, and cloud-native conversions sit alongside the raw archives.
Does Amazon maintain all the datasets it lists?
No, and the registry says so itself: datasets are provided and maintained by a variety of third parties under a variety of licenses, and AWS hosts almost none of the content directly. Each entry names its maintainer - the OpenStreetMap Foundation and Pacific Atlas for the OSM archive - and carries its own license terms.
Which formats do the member datasets use?
Whatever each maintainer chose. The catalog spans S3 object stores, protobuf-based PBF archives, Parquet tables, COG rasters, NetCDF and Zarr arrays. The OSM entry pairs raw PBFs with cloud-native conversions so columnar engines can read derived road geometry without format conversion in between.
Is this catalog useful for highways and rail-track work specifically?
Yes, through its anchor entries rather than a dedicated transport section. The global OSM archive supplies road and rail way geometry worldwide with change feeds for tracking edits over time, while partner collections add climate, imagery and environmental context around corridors. For measured US series, pair it with USDOT and National Bridge Inventory records in the same slice.
Can a sample be cut to specific entries or collections?
Yes. Name the entries, partner collections or domains you care about - the OSM planet spine alone, or a whole NOAA collection - and the sample arrives shaped to that scope with the complete field dictionary attached. Samples precede any commitment, and per-entry terms travel with every extract.
How fresh is the catalog itself?
New entries are contributed continuously and tracked in a public changelog, so the listing grows rather than versioning. Member datasets underneath run from static archives to minutely replication streams. One research note: the inline search reported 13 matches while the page linked 1,069 detail pages, so enumeration beats search for completeness.
Datasets that pair with this one
- Overture Maps Transportation Theme The same roads, re-modelled as an entity graph with consistent IDs - where this catalog gives raw planet files.
- Geofabrik OpenStreetMap Data Extracts Country- and region-sized OSM extracts when the full planet is more than your corridor needs.
- OpenStreetMap Overpass API Query OSM selectively instead of hauling the archive - the surgical alternative to bulk planet pulls.
- open data catalog, defined What separates a metadata directory like this one from the datasets it lists - and why the gap matters.
- geoparquet cloud-native release, defined The columnar format behind the registry's cloud-native conversions, explained.
- Best highways & railtracks datasets Where the registry ranks against ten measured-series records in its slice, scored for 2026.
See the rows before you pay anything.
Name this dataset and we send real records from it — scoped to the fields you asked for.