News result
A query-bound result shown in the Google News vertical.
- Headline and destination
- Source and snippet
- Query and request context
News & Artikel
Discover public news results through the documented Google Search API news vertical, access eligible public article pages, or scope a maintained feed that keeps publication, article, revision, and provenance records distinct.
Story record model
Model discovery, publication, content, people, revisions, and subjects separately so a result snippet never masquerades as a complete article.
A query-bound result shown in the Google News vertical.
The public publisher identity and source page associated with an observation.
Source-linked headline, byline, body, metadata, and media references where the eligible page exposes them.
A successive observation of selected article fields, never an inferred editorial intent.
An enrichment layer produced under agreed models or rules and kept separate from source facts.
Discovery is not the same as full article extraction. A snippet, article page, author byline, publisher, revision, and topic each carry a different evidence boundary.
Coverage contract
Coverage is confirmed by discovery surface, article source, page family, locale, required fields, and responsible-use boundary.
Record schema
Keep what the source exposed separate from what the collection observed and what a downstream enrichment inferred.
{
"record_id": "story_0042",
"object_type": "article_observation",
"discovery": {
"query": "energy outlook",
"surface": "google_news",
"rank": 3
},
"article": {
"headline": "Grid investment accelerates",
"publication": "Example Daily",
"body_state": "observed"
},
"published_at": "YYYY-MM-DDT06:10:00Z",
"observed_at": "YYYY-MM-DDT08:42:00Z",
"source_url": "https://news.example/story/42",
"schema_version": "news.v1"
}published_at is source-attributed publication time; observed_at is collection time. Neither should overwrite the other.
Identity and lineage
Preserve the publisher, URL, discovery context, and article observation before adding canonical URLs, clusters, entities, or topic labels.
Discovery key
Source cues
Content state
Relationship
Freshness semantics
Set collection cadence by workflow, then retain the time attributed by the source alongside request, observation, delivery, and change times.
The displayed publication time, exactly as interpreted under the contract.
The article or result was observed in the requested context.
The record became available to the customer.
The observation entered or left the monitored history.
Quality and missingness
Carry the state that explains why a field is missing so discovery, parsing, source availability, and policy exclusions remain distinguishable.
The record passed agreed discovery or article field checks.
The source page did not show the requested byline, time, or body field.
The page or relationship needs inspection under the contracted rules.
Excluded, restricted, failed, or unavailable remains distinct from an empty article.
Required keys, types, result counts, body-state reason, timestamp shape, and schema version.
URL format, source attribution, publication-time parse state, language, and optional content hash.
Your team approves representative records, source set, permissible use, and decision thresholds; we operate the agreed technical checks.
Operating model
Maximum control
Applications
The record layer supplies source-linked observations; your team defines analytical models, materiality, thresholds, and decisions.
Track named queries, sources, and public story fields over time.
Results · articles · revisionsBuild reviewable corpora for trend, narrative, and source analysis.
Articles · topics · provenancePrepare time-aware, source-linked inputs for approved retrieval workflows.
Text state · source URL · timestampsSurface new public coverage for human review without presenting it as a conclusion.
Discovery · entities · review stateCompare public publishing patterns, sections, and revision behavior.
Publication · article · first seenRetain agreed metadata and permissible content with traceable collection history.
Schema version · provenance · change historyRepresentative pilot
Use real query contexts and representative public article pages, including ordinary stories, revisions, absent fields, changed templates, and restricted states.
Queries, sources, article fields, contexts, cadence, and permitted use.
Results and pages that represent source templates, missing fields, and changes.
Source evidence, times, body states, lineage candidates, and exceptions.
Approve the schema, source set, operating model, and technical thresholds.
Evaluation FAQ
Answers describe documented capabilities and the terms that a source-specific pilot must confirm.
The Google Search API supports the Google News result vertical through the tbm=nws parameter. It is suited to discovering visible news results for a defined query and request context. Eligible public article pages can also be collected through the Scraper API or Browser API, with extraction handled by your workflow or by an agreed delivery scope.
No. Discovery is not the same as full article extraction. A result can expose a headline, destination URL, source, snippet, and displayed publication cue without exposing the complete article body. Article fields require a separate eligible public page and source-specific extraction contract.
A scoped article record can include the source URL, canonical URL when exposed, headline, standfirst, byline, displayed publication time, section, visible body, media references, tags, language, and provenance. Field presence varies by publisher and template, so source-agnostic normalized schemas are pilot first.
published_at records the time attributed to the story by the source when one is exposed. observed_at records when the collection produced the observation. They remain separate because a story may be discovered, revised, or collected long after publication.
A scheduled or managed pilot can evaluate URL canonicalization, publisher identity, headline similarity, entity cues, and time windows to create candidate story clusters. The original articles remain separate, and ambiguous cluster membership is retained for review rather than silently merged.
A recurring scope can retain successive observations, content hashes, first-seen and last-seen times, and selected field changes. Revision history begins when monitoring begins unless a separate historical source is validated, and source corrections are not interpreted beyond what the public page shows.
The record can distinguish a search result with no article collection, a field not exposed by the page, a page unavailable in the requested context, a collection failure, a changed extraction pattern, and an excluded source. These states are not collapsed into an empty article body.
Paywalled, login-only, member-only, and otherwise restricted article bodies are not standard scope. A public discovery result or public metadata page does not authorize collection of restricted content behind it.
No blanket copyrighted redistribution right is implied. Customers remain responsible for source terms, copyright, licensing, retention, access controls, and permitted downstream use. A delivery can be limited to metadata, links, excerpts, or other agreed fields when appropriate.
Proxy customers maintain their own collectors and parsers. WebScrapingAPI maintains access within documented API boundaries. Scheduled feeds and Managed Web Data can include contracted extraction, schema maintenance, quality monitoring, and delivery operations for the agreed source set.
News and articles
Bring the queries, sources, fields, contexts, and use case. We will help separate what is documented from what the representative pilot needs to prove.