Archive search, retrieval, packaging, and delivery.
- Apply the selected archive filters
- Prepare the matched snapshot set
- Keep snapshot metadata with stored pages
- Deliver through the selected archive path
Datensätze & Feeds
Search past captures by domain, URL pattern, date range, language, and content type. Receive matching pages with snapshot metadata for research, backfills, AI data, and change analysis.
Query design
Define what belongs in the search before retrieving pages. A focused brief keeps the matched set relevant, easier to estimate, and simpler to process downstream.
01 · source
Name the public domains and path patterns that contain the pages your research, index, or corpus needs.
02 · time
Use a precise date range so each matched capture belongs to the period you intend to analyze or backfill.
03 · content
Narrow by language and content type when the receiving workflow expects a particular kind of source material.
04 · destination
Choose how the archive package should arrive and keep page content beside its capture metadata from the first handoff.
Historical context
Matched captures create a dated history for each URL or URL family. Keep the stored page and capture context together so downstream teams can compare versions with a clear time axis.
Retain the page as it appeared at the recorded capture time.
Preserve another capture for field, text, link, or layout comparison downstream.
Complete the requested window with the latest matching capture in the result set.
Review the matched scope against the selected domains, paths, dates, languages, and content types. Refine the query until the result represents the historical slice your workflow can use.
Retrieval and delivery
The product journey is deliberate: define filters, review the matched scope and estimate, prepare the selected historical pages, then deliver content and metadata together.
Domains, URL patterns, date range, language, content type, and destination.
Check that the result shape aligns with the historical question before retrieval.
Package stored pages with URL, capture time, language, type, and collection metadata.
Choose S3 or webhook delivery for the receiving data workflow.
Web Archive supplies historical pages and context. Managed Data is the operated path when the finished output must be a maintained structured dataset.
Use cases
Web Archive fits workloads that need earlier page states, a targeted backfill, or a repeatable set of historical captures before analysis begins.
Product fit
Web Archive starts from captures already present in a historical window. Other products fit better when the job begins with a new collection, an available schema, or a fully managed recurring data program.
Produkt | Best starting point | What you receive | Operating fit |
|---|---|---|---|
Webarchiv | A historical domain, URL family, and date window | Stored public pages with snapshot metadata | Historical discovery, backfill, and change analysis |
An available data domain and schema | Prepared dataset details and representative samples | Discover an existing collection before choosing delivery | |
A defined structured collection, cadence, and destination | Recurring scheduled structured deliveries | Keep a known data scope flowing into downstream systems | |
A new bounded multi-page collection | Collection results aligned during evaluation | Collect related eligible pages from approved entry points | |
A supported source target | Maintained structured records on request | Application-led source access through a known schema | |
A custom source, schema, cadence, and destination | A fully operated structured delivery | Hand off collection, extraction, quality, and maintenance |
Pricing orientation
Current pricing is the source of truth. A useful review starts with the query scope, expected matched volume, content package, destination, and support requirements.
Archive evaluation
Bring a domain set, path patterns, date range, language and content-type needs, plus the system that will receive the archive package.
FAQ
Use these answers to shape the first archive query, understand the delivered record, and choose the right product boundary.
Web Archive helps teams find historical public-page snapshots already present in the archive. Define the relevant domains, URL patterns, date range, language, and content type, then receive the matching pages with capture context.
A delivered snapshot pairs the stored page with snapshot metadata such as the source URL, capture time, language, content type, and collection context needed to organize or compare it downstream.
Scope the query with a domain or domain set, URL patterns, a date range, language, and content type. Tighter filters make the matched set easier to review and keep the delivery focused on the historical slice your workflow needs.
Yes. The evaluation flow is designed around reviewing the matched scope and estimate before the snapshot package is prepared. You can refine filters when the initial result set is broader or narrower than the intended research question.
The archive package can be delivered to S3 or webhook, with stored pages and their snapshot metadata kept together. Choose the path that best fits your ingestion, storage, and processing workflow.
Web Archive searches for historical captures already present across a chosen time window. Crawl API is the better fit when you need to plan a new bounded collection across related eligible public pages.
Web Archive delivers historical pages with capture metadata. Datenmarktplatz helps teams discover and sample prepared datasets, Datenfeeds provides recurring scheduled structured delivery, and Data API returns maintained structured records for supported source targets.
Yes. Retrieve captures for the same URL or URL family across a defined date range, retain each capture time, and compare the versions in your downstream research, indexing, monitoring, or extraction workflow.
Bring the domains, URL patterns, date range, language and content-type filters, expected matched scope, delivery preference, and support requirements. Current pricing and a workload review remain the source of truth for commercial details.
Historical data brief
Share the domains, paths, dates, content filters, and destination. We will help shape the query and the right delivery path.