Recruitment & HR
See the labor market beyond your ATS.
Turn public job postings and employer career activity into source-linked records for workforce planning, recruiting strategy, compensation research, labor-market analysis, and data products—without building a candidate database.
Public job and company signals only. Candidate profiles, applications, personal contact data, private employee systems, and automated hiring decisions are not part of the standard scope.
01
Define the labor market
Employers, occupations, public sources, geographies, and cadence.
02
Inspect the lifecycle
New, changed, duplicated, reposted, closed, missing, and failed states.
03
Choose the operating boundary
Infrastructure, API access, scheduled feed, or managed program.
Decision coverage
Talent teams need posting history before labor-market interpretation.
Public postings reveal observable hiring demand. They do not establish completed hires, internal headcount, payroll, applicant quality, employee performance, or workforce plans.
01
Workforce planning
See where external hiring demand changes.
Which functions, occupations, locations, and seniority levels are expanding, contracting, or shifting across the defined employer market?
Active posting panel
First and last seen
Function, level and location
02
Talent acquisition
Map competition for the roles you need.
Which employers are publicly recruiting for comparable roles, skills, work models, and markets?
Normalized role
Published skills
Employer, location and work mode
03
Compensation & rewards
Compare the pay employers publish.
What salary ranges, currencies, periods, benefits, and employment terms are publicly displayed for comparable roles?
Disclosed range only
Currency and pay period
Role and location context
04
Employer brand & people strategy
Read the public employee proposition in context.
How are employers describing responsibilities, benefits, flexibility, culture, and career paths across public hiring surfaces?
Description and benefits
Work-model language
Career-page changes
05
Labor-market research
Measure observable demand with its denominator.
How is the approved posting panel changing by occupation, industry, geography, employer, and time?
Deduplicated posting counts
Coverage by source
Lifecycle and collection states
06
Talent products & data engineering
Build on records users can trace.
Can every listing, alert, trend, and enrichment field be traced to its public source and collection state?
Canonical source URL
Schema and timestamps
Update, removal and duplicate states
Public-source coverage
Define the public hiring market you need to observe.
Coverage is an approved employer, source, page-type, role, geography, field, and cadence brief—not an assumed view of every vacancy or the entire labor market.
Source families
Public Employer career sites Search, detail, location, team, and public ATS pages
Public Job boards & aggregators Mainstream listings and public discovery surfaces
Review Niche & public portals Industry, association, government, and regional boards
Review Employer context Company sites, directories, and public workplace information
Approved labor-market brief
Record context
01 Employer & posting Source IDs, official domain, canonical URL
02 Role & workplace Title, function, seniority, location, work mode
03 Public offer Employment type, disclosed pay, benefits
04 Lifecycle & evidence First seen, changed, closed, missing, failed
Scope states
Dokumentiert
General public-page access, browser rendering, search, proxies, and documented API behavior.
Pilot first
Source-specific fields, salary coverage, taxonomy mapping, employer identity, duplicates, removals, and history.
Nicht standardmäßig
Candidate data, CVs, personal contacts, applications, login-gated portals, private ATS or HRIS records, payroll, or inferred protected traits.
Public availability does not remove privacy, employment-law, source-term, copyright, retention, or purpose review.
Inspectable data contract
A job record should preserve both meaning and lifecycle.
Keep the employer, source posting, raw and normalized role fields, public offer, observation time, duplicate state, and lifecycle together so every trend remains reviewable.
01 · Identity
Which employer and posting?
Source job ID, canonical URL, employer name, official domain, source employer ID, and contracted company reference.
02 · Role & place
What work is being advertised?
Raw and normalized title, function, occupation, seniority, department, location, country, and work mode.
03 · Public offer
What did the source disclose?
Employment and contract type, published compensation, currency, period, benefits, skills, qualifications, and description.
04 · Lifecycle & evidence
What happened to the record?
Published, updated, expiry, first-seen, last-seen, last-changed, lifecycle, duplicate, missing, source URL, capture, and schema states.
Illustrative job record Not customer data
- job_ref
job-demo-042
- employer_ref
company-demo-17
- source_job_id
ATS-8431
- title_raw
Senior Data Engineer
- title_normalized
Data Engineer
- work_mode
hybrid
- compensation_state
not_published
- first_seen
2026-07-14
- last_seen
2026-07-23
- lifecycle_state
updated
- schema_version
jobs.v1
Source linked Raw title retained No pay inferred
Employer identity, deduplication & lifecycle
Comparable labor data starts with posting identity—and ends with lifecycle.
Anchor the employer and source posting before normalizing role language. Preserve ambiguity when two URLs may represent a duplicate or repost.
01
Anchor the employer
Use official domain, source employer ID, company name, and location before fuzzy company labels.
02
Prefer exact posting keys
Use source job ID and canonical URL before title, location, description, and timing fingerprints.
03
Keep raw and normalized values
Retain source language beside role, occupation, seniority, skill, and work-mode taxonomies.
04
Emit explicit lifecycle states
New, active, changed, duplicate candidate, repost candidate, source closed, not observed, failed, or excluded.
Cross-source employer resolution, posting deduplication, taxonomy mapping, and lifecycle history are separately scoped feed or managed capabilities. They are not implied for every self-service page response.
Posting identity desk Review state
Employer anchor
Northstar Labs · official careers
Fictional entity · source employer ID retained
Exact observation
ATS-8431 · canonical job URL
Source ID and employer agree
Board observation
Exact duplicate
Same source ID · same role · same location
Separate board URL
Possible repost—review
Title matches · timing differs · no forced merge
Quality & inference boundary
A missing page is not a closed role. A new URL is not always a new job.
Keep source-declared closure, not-observed results, collection failures, duplicates, reposts, and content changes separate so the feed never invents a hiring event.
Posting history Job demo 042
4 observations
14 Jul · source Posting first observed Active
16 Jul · board Exact duplicate found Grouped
21 Jul · source Work-mode text changed
23 Jul · collector Detail page retrieval failed
Observed
Required fields evaluated
The public posting returned and the contracted fields were processed.
Source closed
The source declared closure
A closed or expired label is retained with its source and observation time.
Nicht beobachtet
No matching record appeared
This is not proof of a closed requisition, completed hire, or reduced demand.
Failed
Collection did not complete
No labor-market or lifecycle state follows from a failed request.
WebScrapingAPI observes
Public job and company signals
Posting content, disclosed terms, employer context, source identifiers, lifecycle evidence, and collection states.
Contracted processing adds
Structure and continuity
Employer matching, normalization, duplicate review, lifecycle history, quality checks, and delivery when specified.
Your team determines
Every people decision
Workforce planning, recruiting strategy, compensation policy, candidate assessment, hiring decision, employment action, and lawful use.
Four operating models
Choose how public job evidence enters your talent workflow.
Each model separates WSA-operated collection and delivery from your purpose, privacy, employment-law review, recruiting methods, and every candidate or hiring decision.
Infrastructure for your collectors
Run your own job-market collection through proxy infrastructure.
WebScrapingAPI operates contracted proxy-network features. Your team owns approved sources, collectors, extraction, employer identity, normalization, duplicates, lifecycle, schedules, quality, history, storage, and decisions.
Proxy infrastructure ownership for recruitment data
Lifecycle responsibility Owner
Purpose, source rights, privacy & exclusions Ihr Team von Experten
Proxy routing, rotation & contracted location options WSA
Collectors, rendering, extraction & schema Ihr Team von Experten
Identity, taxonomy, lifecycle, quality & delivery Ihr Team von Experten
Recruiting, workforce & hiring decisions Ihr Team von Experten
On-demand public-page access
Call approved public job pages without operating the access layer.
WebScrapingAPI maintains documented request access, retries, supported rendering, and the chosen endpoint response. Your team owns raw-page parsing, employer mapping, normalization, lifecycle, and downstream use.
Web access API ownership for recruitment data
Lifecycle responsibility Owner
Purpose, employers, approved sources & request brief Ihr Team von Experten
API access, routing, retries & supported rendering WSA
Extraction & schema returned by the endpoint By endpoint
Employer identity, taxonomy, duplicates, history & quality Ihr Team von Experten
Recruiting, workforce & hiring decisions Ihr Team von Experten
Recurring structured delivery
Receive agreed job-posting records on a defined cadence.
WebScrapingAPI operates the contracted collection, extraction, schedule, schema checks, source maintenance, quality, and delivery. Employer matching, normalization, deduplication, and lifecycle are included only when specified.
Scheduled recruitment feed ownership
Lifecycle responsibility Owner
Purpose, employer panel, sources & acceptance rules Your team + WSA
Access, collection & extraction WSA when contracted
Employer identity, taxonomy, duplicates & lifecycle WSA when contracted
Scheduling, quality, maintenance & delivery WSA
Recruiting, workforce & hiring decisions Ihr Team von Experten
Operated labor-data program
Hand off the maintained public job-data operation.
Bring the labor-market purpose, employers, roles, sources, fields, taxonomy, cadence, privacy requirements, and destination. WebScrapingAPI designs and operates the agreed workflow with your team.
Managed recruitment data program ownership
Lifecycle responsibility Owner
Purpose, privacy, source rights & exclusions Ihr Team von Experten
Source onboarding, collection & extraction WSA when contracted
Employer identity, normalization, duplicates & lifecycle WSA when contracted
Scheduling, quality, maintenance, exceptions & delivery WSA
Recruiting, workforce & hiring decisions Ihr Team von Experten
Representative labor-data pilot
Prove the posting lifecycle with records your teams can inspect.
Start with one role family, employer panel, and geography. Include normal listings, duplicates, reposts, changes, missing pay, source closures, not-observed records, and collection failures.
- 01 · Frame
Define the labor market
Choose employers, occupations, geographies, public sources, fields, taxonomy, cadence, exclusions, and destination.
- 02 · Sample
Collect representative states
Include multi-location roles, duplicate URLs, repost candidates, content changes, missing fields, closures, and failures.
- 03 · Validate
Agree the lifecycle contract
Review employer identity, role mapping, duplicate rules, state definitions, history, quality checks, and acceptance.
- 04 · Operate
Launch the right handoff
Assign collection and maintenance ownership, connect delivery, monitor source continuity, and preserve privacy controls.
A pilot validates collection and the data contract—not a recruiting strategy, labor forecast, candidate assessment, or hiring decision.
Evaluation questions
What recruitment and labor-data teams should confirm before collection.
Sources, candidate-data boundaries, duplicates, closure states, pay, taxonomy, history, cadence, maintenance, and governance—answered directly.
Which recruitment sources can be covered?
Eligible public employer career pages, job boards, aggregators, niche portals, and related public company sources can be evaluated. Exact page types, fields, markets, formats, and cadence are confirmed with representative requests.
Does WebScrapingAPI provide candidate profiles, CVs, or personal contact data?
This offering does not provide candidate profiles, employee profiles, CVs, applications, personal contact data, private people systems, or person-level assessment. The scope centers public job postings and company signals.
How are duplicate listings and reposts handled?
A contracted workflow can use source IDs, canonical URLs, employer identity, title, location, description fingerprints, and lifecycle timing. Ambiguous observations remain candidate duplicates or repost candidates rather than being merged silently.
How do you know when a job has closed?
A source-declared closed or expired state is distinct from a listing that was merely not observed. Collection failures are separate again, preventing gaps from becoming false closure events or implied hires.
Is compensation available for every posting?
No. Compensation is delivered only when publicly displayed and in scope. Currency, pay period, location, role, and source context remain attached. Missing pay is not automatically predicted or inferred.
Can titles, skills, and occupations be normalized?
Yes when included in a scheduled or managed contract. Raw source values are retained beside the normalized taxonomy, mapping state, and taxonomy version so users can review every transformation.
Can historical job-posting data be delivered?
Forward history can begin when recurring collection starts. Backfill depends on eligible public source history or archives and is separately validated for coverage, identifiers, and schema consistency rather than assumed.
How fresh can the records be?
Web access APIs return observations when called. Scheduled and managed programs use a cadence agreed by source, field, and business need. Each record should retain its capture time; no universal real-time frequency is promised.
Who maintains collection when a source changes?
Proxy and raw page-access customers maintain their collectors or parsers. WebScrapingAPI maintains the documented API layer and maintains contracted connectors, extraction, quality, source-change work, and delivery for scheduled and managed programs.
How are records delivered and governed?
Delivery can use APIs or contracted structured destinations. Production scope defines source eligibility, candidate and personal-data exclusions, fields, retention, access, taxonomy, quality states, permitted use, and ownership of every workforce or hiring decision.
Related paths
Continue with the closest talent data path.
Build your public labor-data foundation
Validate employer identity, posting state, duplicates, and repost rules.
Share the employers, roles, public sources, geographies, taxonomy, cadence, history, exclusions, and destination. We’ll map supportable coverage and produce representative records including duplicates, removals, missing fields, and collection gaps.