Zum Inhalt springen

Datensätze & Feeds

Web Archive for historical public-page snapshots.

Search past captures by domain, URL pattern, date range, language, and content type. Receive matching pages with snapshot metadata for research, backfills, AI data, and change analysis.

  • Query scope Domain and URL pattern
  • Time context Date range and capture time
  • Snapshot package Stored page with metadata
  • Delivery path S3 or webhook

Query design

Start with the smallest historical slice that answers the question.

Define what belongs in the search before retrieving pages. A focused brief keeps the matched set relevant, easier to estimate, and simpler to process downstream.

  1. 01 · source

    Domain · URL pattern

    Choose domains and URL families.

    Name the public domains and path patterns that contain the pages your research, index, or corpus needs.

  2. 02 · time

    Start date · end date

    Set the historical window.

    Use a precise date range so each matched capture belongs to the period you intend to analyze or backfill.

  3. 03 · content

    Language · content type

    Keep the relevant page types.

    Narrow by language and content type when the receiving workflow expects a particular kind of source material.

  4. 04 · destination

    S3 · webhook

    Design for the receiving pipeline.

    Choose how the archive package should arrive and keep page content beside its capture metadata from the first handoff.

Historical context

Read a page as a sequence, not a single current state.

Matched captures create a dated history for each URL or URL family. Keep the stored page and capture context together so downstream teams can compare versions with a clear time axis.

Illustrative snapshot historyOne URL · three captures
  1. Earlier
    Baseline capture

    Retain the page as it appeared at the recorded capture time.

    snapshot 01
  2. Revision
    Changed page version

    Preserve another capture for field, text, link, or layout comparison downstream.

    snapshot 02
  3. Later
    Later archived state

    Complete the requested window with the latest matching capture in the result set.

    snapshot 03
Make coverage inspectable

Review the matched scope against the selected domains, paths, dates, languages, and content types. Refine the query until the result represents the historical slice your workflow can use.

Retrieval and delivery

Move from a scoped match set to a usable archive package.

The product journey is deliberate: define filters, review the matched scope and estimate, prepare the selected historical pages, then deliver content and metadata together.

  1. 01 · define

    Submit the query brief.

    Domains, URL patterns, date range, language, content type, and destination.

  2. 02 · review

    Inspect the matched scope.

    Check that the result shape aligns with the historical question before retrieval.

  3. 03 · prepare

    Keep snapshots and context together.

    Package stored pages with URL, capture time, language, type, and collection metadata.

  4. 04 · deliver

    Send the package downstream.

    Choose S3 or webhook delivery for the receiving data workflow.

WebScrapingAPI operates

Archive search, retrieval, packaging, and delivery.

  • Apply the selected archive filters
  • Prepare the matched snapshot set
  • Keep snapshot metadata with stored pages
  • Deliver through the selected archive path
Your team defines

The question, usable scope, and downstream data product.

  • Approved domains, paths, and historical window
  • Language and content requirements
  • Parsing, comparison, indexing, or model workflow
  • Storage, retention, access, and downstream decisions
Need structured output?

Move extraction, normalization, quality monitoring, and recurring delivery into Managed Data.

Web Archive supplies historical pages and context. Managed Data is the operated path when the finished output must be a maintained structured dataset.

Explore Managed Data

Product fit

Choose history, live collection, structured records, or operated delivery.

Web Archive starts from captures already present in a historical window. Other products fit better when the job begins with a new collection, an available schema, or a fully managed recurring data program.

Produkt

Best starting point

What you receive

Operating fit

Webarchiv

A historical domain, URL family, and date window

Stored public pages with snapshot metadata

Historical discovery, backfill, and change analysis

Datenmarktplatz

An available data domain and schema

Prepared dataset details and representative samples

Discover an existing collection before choosing delivery

Datenfeeds

A defined structured collection, cadence, and destination

Recurring scheduled structured deliveries

Keep a known data scope flowing into downstream systems

Crawl-API

A new bounded multi-page collection

Collection results aligned during evaluation

Collect related eligible pages from approved entry points

Data-API

A supported source target

Maintained structured records on request

Application-led source access through a known schema

Verwaltete Daten

A custom source, schema, cadence, and destination

A fully operated structured delivery

Hand off collection, extraction, quality, and maintenance

Pricing orientation

Price the historical slice your team will actually use.

Current pricing is the source of truth. A useful review starts with the query scope, expected matched volume, content package, destination, and support requirements.

FAQ

Web Archive questions for a focused evaluation.

Use these answers to shape the first archive query, understand the delivered record, and choose the right product boundary.

Scope an archive query

What is Web Archive?

Web Archive helps teams find historical public-page snapshots already present in the archive. Define the relevant domains, URL patterns, date range, language, and content type, then receive the matching pages with capture context.

What does an archive snapshot include?

A delivered snapshot pairs the stored page with snapshot metadata such as the source URL, capture time, language, content type, and collection context needed to organize or compare it downstream.

How can I narrow an archive query?

Scope the query with a domain or domain set, URL patterns, a date range, language, and content type. Tighter filters make the matched set easier to review and keep the delivery focused on the historical slice your workflow needs.

Can I review the matched scope before delivery?

Yes. The evaluation flow is designed around reviewing the matched scope and estimate before the snapshot package is prepared. You can refine filters when the initial result set is broader or narrower than the intended research question.

How is archive data delivered?

The archive package can be delivered to S3 or webhook, with stored pages and their snapshot metadata kept together. Choose the path that best fits your ingestion, storage, and processing workflow.

How is Web Archive different from Crawl API?

Web Archive searches for historical captures already present across a chosen time window. Crawl API is the better fit when you need to plan a new bounded collection across related eligible public pages.

How is Web Archive different from Data Marketplace, Data Feeds, and Data API?

Web Archive delivers historical pages with capture metadata. Datenmarktplatz helps teams discover and sample prepared datasets, Datenfeeds provides recurring scheduled structured delivery, and Data API returns maintained structured records for supported source targets.

Can Web Archive support change-history analysis?

Yes. Retrieve captures for the same URL or URL family across a defined date range, retain each capture time, and compare the versions in your downstream research, indexing, monitoring, or extraction workflow.

What should I provide for pricing?

Bring the domains, URL patterns, date range, language and content-type filters, expected matched scope, delivery preference, and support requirements. Current pricing and a workload review remain the source of truth for commercial details.

Historical data brief

Turn a historical question into a focused archive package.

Share the domains, paths, dates, content filters, and destination. We will help shape the query and the right delivery path.