Best Job Posting Data Providers in 2026: Datasets, APIs, and Custom Collection Compared

The best job posting data provider is the one whose sources, fields, geography, and refresh cadence match what you actually need. Record count is the wrong first filter. Raw posting volume gets inflated by duplication, because the same job is syndicated across boards, and by listings that stay live after the role is gone. A dataset advertising billions of records can still miss the twelve regional boards your analysis depends on.
For labor market and workforce research, Lightcast, LinkUp, and Revelio Labs are the strongest packaged options. For API access, go-to-market signals, and high-volume feeds, look at Coresignal, TheirStack, PredictLeads, Canaria, Signalbase, and Xverum. Adzuna and Techmap serve developers and cost-sensitive buyers. People Data Labs has a job postings product, still labeled beta as of its July 2026 release.
At Ficstar, a fully managed web scraping and data collection company, we collect job listings at scale, millions a week across hundreds of sources. That work means we spend a lot of time looking at exactly where packaged datasets stop. When no product on the market covers the sources, fields, or cadence a team needs, the alternative is collecting the data from the source. That is the second half of this guide.
All provider figures below were checked in September 2026 against each provider's own documentation, pricing, or product pages. Record counts, archive start dates, and refresh cadences change often in this market, so verify them again before you sign anything.
How we evaluated job posting data providers
We applied the same eleven criteria to every provider. These are the questions that decide whether a dataset is usable for your specific job, and they are worth asking on every vendor call.
Sources covered. General job boards, company career pages, ATS platforms, professional networks, staffing agency sites, government employment portals, university career centers, and regional boards. Ask for the actual source list, not a count.
Refresh cadence and closure detection. How often each posting is revisited, and how the provider decides a listing is closed.
Deduplication. Whether the same role appearing on five boards resolves to one record, and what method does the resolving.
Normalization. How job titles, locations, salary formats, seniority, and company names are standardized across inconsistent sources.
Historical depth. The archive start date, and whether history is available on day one or accrues from your start date.
Enrichment. Salary parsing, skills taxonomies, occupation codes, company firmographics, and seniority inference.
Delivery. API, bulk files, webhooks, S3, SFTP, and which formats are supported.
Geographic coverage. Countries covered, and how evenly. Global coverage is often thin outside a handful of markets.
Compliance posture. What the provider collects, what it refuses to collect, and how it handles site terms.
Pricing model. Per record, per credit, flat license, or annual contract, and how the cost curve behaves as volume grows.
Sample or trial. Whether you can evaluate real records before committing, and on what terms.
Why record count is the wrong first filter
Duplication is the main reason posting counts overstate hiring demand. Research presented at ACM WI-IAT 2021 on duplicate detection in online job postings cites figures from jobs-data vendor Textkernel showing that an average job ad is reposted two to five times depending on the country. At that rate, duplicates account for as much as 50 to 80 percent of crawled postings. TheirStack describes the same pattern in its own documentation, with jobs appearing three to five times across platforms.

Ghost postings are the second distortion. Greenhouse's review of its own platform found that 18 to 22 percent of job posts were ghost listings, as reported by MarketBeat in January 2025. A 2025 employer survey by Clarify Capital found nearly one in three employers had postings active for more than 30 days. The Congressional Research Service, in its April 2025 report on ghost job postings, notes there are no official statistics on how common they are, and points to two mundane causes: third-party boards auto-copying listings, and employers not removing roles that have been filled.
Even clean posting data measures intent rather than hiring. The gap between US job openings and hires has held at roughly 30 percent every month since 2021, according to a MyPerfectResume analysis of JOLTS data reported by HR Dive in November 2025, with the widest gaps in government, education and health, information, and financial activities. The Bureau of Labor Statistics defines a job opening in JOLTS narrowly: a specific position with work available, able to start within 30 days, under active external recruiting. Its July 2026 release counted 7.3 million openings against 5.1 million hires. Statistics Canada's Job Vacancy and Wage Survey applies a similar definition for Canadian vacancies. Both are the yardsticks your posting counts should be sanity-checked against.

None of this makes posting data unreliable. It does mean volume claims should be read as raw collection, and that the dedup and closure-detection methods behind them are what you are actually buying.
Best job posting data providers in 2026 compared
The table below compares the packaged datasets and APIs, with Ficstar included as the fully managed alternative for teams whose sources, fields, or cadence fall outside any off-the-shelf product. Read the packaged providers on record count, history, and refresh; read Ficstar on what it will collect to your specification, since it builds to your requirement rather than selling a fixed dataset.
Provider | Built for | Sourcing and coverage (provider-stated) | History from | Refresh | Trial or sample |
Coresignal | Deduplicated cross-board volume via API or dataset | 475M+ postings, 85+ fields, multi-source and raw single-source products | 2020 | Active postings revisited within 24 hours | 7-day free trial, self-service samples |
Lightcast | Workforce, education, and government analytics | 220,000+ websites worldwide including 100+ government sources | 2010 (US postings) | Daily | Contact sales |
LinkUp | Employer-sourced research and investment signals | 350M+ postings indexed from employer career sites | 2007 | Daily | Contact sales, academic access via Dewey Data |
Revelio Labs | Workforce intelligence and economic research | COSMOS: 5B+ postings from employer sites, boards, and staffing firms | 2021 (COSMOS) | Daily | Discounted academic delivery |
Canaria | Research-grade deduplicated postings via marketplaces | 1B+ deduplicated postings, 100+ enriched fields | Not published | Hourly on US records per marketplace listing | Free sample |
Xverum | Structured B2B feeds tied to company data | 468M+ postings per marketplace profile, 10M+ analyzed daily | Not published | As fast as 24 hours | Free samples |
PredictLeads | Hiring signals inside company intelligence | 270M+ records across 2.7M companies, from company sites and ATS | 2016 | Every 36 hours | Contact sales |
TheirStack | Sales intelligence and technographics | 356,000+ websites across 195+ countries | 2021 | 73% same day, 90% within 24 hours | Free tier, paid from about $59/month |
Signalbase | Real-time hiring signals for outreach timing | Job boards, careers pages, and social posts across 190+ countries | No historical corpus | Sub-minute | Not published |
People Data Labs | Career-page postings, product in beta | 400,000+ company career pages, 100,000+ companies with active posts | Beta, schema evolving | Daily delivery option in beta | 5,000 free beta credits |
Adzuna | Developers and lightweight applications | Roughly 20 markets, live ads plus salary histograms | Not applicable | Live | Instant free API key, capped daily calls |
Techmap | Low-cost global volume | 8.2M new postings monthly from 127+ sources across 250 countries | 2020 | Ongoing feeds | Free tier, roughly $1 per 1,000 postings |
Ficstar | Teams whose sources, fields, or cadence fall outside any packaged dataset | Custom collection from the public sources you specify, including sites that block or cap scrapers such as Indeed and LinkedIn Jobs; millions of listings a week, managed end to end (a service, not a fixed dataset) | Forward-looking from your start date; historical backfill where a source allows | Any cadence you set: real-time, daily, weekly, or custom | Free trial that is real collection from your own sources, typically two weeks |
Large aggregated job posting datasets and APIs
These providers sell breadth. They pull from many source types and resolve the overlap, which suits teams that want one feed covering as much of the market as possible.
Coresignal
Coresignal sells datasets, APIs, and a no-code dashboard drawing from a single refreshed database, with two job products. Multi-Source Jobs runs postings through an entity-resolution engine to deduplicate across boards. Base Jobs delivers raw single-source records for teams that want to do their own resolution. Its job postings page states 475M+ postings across 85+ data fields, with every active posting revisited within 24 hours and daily discovery of roughly 1.3 million active jobs. Normalization covers titles, industries, locations, seniority, salary ranges, and company matching. The archive begins in August 2020, so teams needing pre-2020 history will need a second source.
Xverum
Xverum sells structured B2B datasets across people, company, jobs, and places. Its jobs feed analyzes 10M+ postings daily and refreshes as fast as 24 hours, delivered in JSON and Parquet, and its marketplace profile cites 468M+ job postings. The strongest fit is teams that already want company data and want postings tied to the same company records rather than standing alone. Xverum does not publish an archive start date, so historical depth is a question for the sales call rather than something you can confirm from the product page.
People Data Labs
People Data Labs sources postings directly from company-hosted career pages, and its job posting product is labeled Beta on both the product site and the docs. Its July 2026 v35.0 release expanded coverage to 400,000+ unique company career pages, up from 86,000, and to 100,000+ companies with active posts, up from 60,000. The release also reported a description fill rate above 99 percent for active posts and improved deactivation-date accuracy. Beta access includes 5,000 free credits. The documentation states the schema may continue to evolve, which makes this a better fit for teams that can absorb field changes than for a production pipeline that needs a frozen contract.
Employer-sourced and analytics-first providers
These three go to employer career sites first and build analytics on top. That sourcing choice reduces duplication and ghost noise at the collection stage, which is why they dominate research, government, and investment use.

LinkUp
LinkUp indexes postings directly from employer career sites rather than aggregators, and sells datasets, custom feeds, and its Compass analytics tool. Its flagship dataset is described as 350M+ global postings indexed daily since 2007, and a Dewey Data academic listing puts collection at 80,000+ employer career sites with raw records back to 2007. S&P builds the S&P 500 LinkUp Jobs Indices on the data, which says something about how seriously the financial research market treats it. Records carry title, description, location, URL, occupation and sector codes, company identifier, and public ticker where applicable. The tradeoff of career-site-only sourcing is that roles posted exclusively to boards or staffing sites will not appear.
Lightcast
Lightcast is a labor market analytics platform built for workforce boards, education institutions, and government users, and it has the deepest US history in this group. Its documentation puts sourcing at 220,000+ websites worldwide, including company career sites, national and local boards, aggregators, and 100+ government sources, with US postings history back to January 2010 and daily updates. Deduplication compares normalized title, company, and location across a 60-day window. Enrichment maps to O*NET-SOC occupations, a skills taxonomy, credentials, salary, and industry. Independent commentary from Digit Research notes that Lightcast data is scraped, can carry errors, and has diverged from other labor indicators since around 2012, a fair reminder that any postings-based series needs benchmarking against official statistics.
Revelio Labs
Revelio Labs sells workforce intelligence through its Terminal product, raw datasets, and a discounted academic program. Its COSMOS dataset covers 5B+ postings from employer websites, major job boards, and staffing firm boards, deduplicated and unified, updated daily with global coverage. Enrichment includes a role taxonomy, seniority, geography, skills, salary, expected hires, and time-to-fill, with occupation mapping to BLS codes in the US and ILO codes internationally. Revelio also publishes public labor statistics benchmarked against BLS releases, which makes methodology comparison easier than it usually is in this market. COSMOS history begins in 2021, so long-run posting series need a different source even though Revelio's broader workforce data reaches further back.
Hiring-signal and company intelligence providers
These providers treat a job posting as a signal about a company rather than as a record in a labor market corpus. If your output is a prospect list or an account alert rather than an analysis, start here.
PredictLeads
PredictLeads sells company intelligence datasets, and its Job Openings dataset is delivered by API, flat files, webhooks, and MCP. It sources from company websites, career subpages, and ATS integrations. The current product page states 270M+ historical records since 2016 across 2.7M companies, with an average of 9.8 million active jobs at any time and each opening refreshed every 36 hours. All jobs are categorized with O*NET codes, and fields include title, URL, first-seen and last-seen dates, location, category, seniority, description, salary, and contract type. Its docs pages carry conflicting historical figures, so confirm the count and start year with the vendor.
TheirStack
TheirStack combines a job postings API with a technographics database, aimed at sales intelligence teams and job boards. It aggregates from 356,000+ websites including career sites, ATSs, and job boards, with data since 2021 across 195+ countries. Freshness is published rather than implied: 73 percent of new tech postings land the same day, and 90 percent within 24 hours. Enrichment covers salary parsing with min, max, currency, and period, plus seniority, location, company firmographics, and tech stack, with descriptions normalized to Markdown. Pricing is public, starting with a free tier of 50 company credits and 200 API credits a month, and paid plans from about $59 a month with credits rolling over for 12 months.
Canaria
Canaria sells research-grade job market datasets through data marketplaces and direct delivery. Its own site claims 1B+ deduplicated postings with 100+ enriched fields, SOC codes, salary, 40,000+ skills, and one canonical job entity per hiring intent, with sources including major boards. AI and NLP enrichment adds predicted salary, seniority, and normalized titles, and delivery runs through marketplaces, S3, SFTP, and email. A free sample is offered. Worth noting before you build a plan around the numbers: Canaria's figures are self-reported and differ between its own site, which cites 1B+ globally, and its marketplace listing, which cites 800M+ deduplicated US records updated hourly. Ask which figure applies to the product you are buying.
Signalbase
Signalbase is shaped differently from the rest of this section. It is a real-time hiring-signal and outreach tool rather than a research corpus, reading job boards, careers pages, and social posts with sub-minute freshness and delivering through API, webhook, and MCP with firmographic filtering across 190+ countries. There is no bulk historical corpus and no candidate data. If your use case is prospecting timing, that design is the point. If your use case is a labor market study, look elsewhere in this list.
Lower-cost and developer-tier job posting APIs
Not every project needs an enterprise license. Two options serve smaller budgets and prototype work.
Adzuna
Adzuna runs a consumer job search engine and a developer API aggregating listings across roughly 20 markets including the UK, US, Germany, France, India, Australia, Canada, and Brazil. The API returns live ads with title, company, location, salary, description, and redirect URL, plus salary histograms, top-company lists, regional statistics, and historical trend data. The free tier issues an instant App ID and key with a capped number of calls per day, with higher volume available on request. It suits developers building a feature on top of job data more than analysts building a corpus.
Techmap
Techmap, operating as jobdatafeeds.com, sells international job datafeeds, datasets, and an API from a small, cost-focused operation. It reports about 8.2 million new global postings per month from 127+ sources across 250 countries, with roughly 2 million a month for the US and history from January 2020. Fields include source, country, title, URL, text and HTML, salary, location, company, and dates, delivered in JSON by default with other formats on request. Pricing is roughly $1 per 1,000 postings, with historical datasets from around $2,400 per country, plus a free API tier and a free sample dataset. At this price point you are buying volume rather than deep entity resolution.
When to collect job posting data from the source instead
Every provider above sells a fixed product. That works when the product's shape matches your requirement. When it does not, the alternative is collecting the data yourself from the sources you specify, which is what we do at Ficstar as a fully managed service. We are not a dataset provider and we do not license a job postings database. Clients tell us which sources and fields they need, and our job scraping service collects exactly that.
In practice, teams come to us for five reasons.
Sources no vendor covers. Regional boards, industry-specific sites, university career centers, staffing agency sites, and government employment portals that fall outside a packaged product's source list. We collect from general job boards, professional networks, industry sites, company career pages, staffing agencies, government portals, university career centers, and regional boards, at millions of listings a week.
Sites that block or throttle collection. Access is the practical bottleneck in this market. We maintain reliable access to sites that actively prevent scraping, including Indeed and LinkedIn Jobs, using residential proxy networks, CAPTCHA handling, bot-detection avoidance, and rate-limit management.
Per-page result caps. Some job sites display as few as 400 results when thousands exist. Our crawlers are built to page through those limits and retrieve the full set rather than the visible slice.
Fields no vendor extracts. We deliver structured records covering title, company, location including remote, hybrid, or onsite, salary where disclosed, experience and education requirements, required and preferred skills, description, application deadline, posting date, source URL, employment type, and benefits. When a client needs a field outside that set, we add it rather than tell them it is not in the schema.
A cadence and format that fit an existing system. Daily, weekly, real-time, or custom schedules, in CSV, JSON, XML, Excel, or TSV, delivered by API, SFTP, AWS S3, direct database updates, or straight into an ERP, CRM, or BI system. We also pre-filter before delivery by region, industry, job category, experience level, employment type, salary range, or required skills, so teams receive the subset they need rather than everything.
Normalization and deduplication carry over from the packaged world. Titles get resolved across variants like Software Engineer, Software Developer, and SWE. Locations and salary formats are standardized across annual, hourly, currency, and range differences. The same posting appearing on a company site, Indeed, and LinkedIn collapses into one record.

We have delivered job listings collection for government agencies analyzing local job availability, in-demand skills, and public-sector salary benchmarks, and for recruitment analytics companies using postings as their core product input. Workopolis is among the clients on our public roster. At the company level, we have been operating since 2005, serve 200+ enterprise customers, and hold a 5.0 out of 5.0 rating on G2 across 61 reviews as of April 2026. Every project is backed by a 100 percent satisfaction guarantee.
One honest caveat. Custom collection does not eliminate the maintenance and terms-of-service burden that comes with collecting data from the web. Sites change layouts, boards add defenses, and crawlers need upkeep. A fully managed service absorbs that work rather than making it disappear.
Should you license a job postings dataset or collect your own?
Most teams can answer this in one pass by checking which column they sit in.
Signal | License a packaged dataset | Collect from the source |
Source coverage | A vendor already covers the boards and career sites you care about | You need specific sources no vendor lists, including blocked or result-capped sites |
Fields | The published schema contains what you need | You need fields no vendor extracts |
History | You need deep archive on day one and cannot wait to accrue it | Forward-looking collection is enough for your analysis |
Enrichment | You want skills taxonomies, occupation codes, or salary models you would not build | You do your own enrichment downstream |
Cadence and format | Standard daily or weekly delivery fits your pipeline | You need a defined cadence and a format that fits an existing system |
Exclusivity | A shared dataset is fine | The data needs to be yours rather than sold to everyone in your category |
Maintenance | You want a vendor's roadmap, not a collection project | You want collection maintained against your source list as sites change |
Volume | Modest and predictable | Large, growing, or spread across many niche sources |
You can also do both. Plenty of teams license a broad dataset for baseline coverage and run custom web scraping for the sources, fields, or geographies that dataset misses. Ficstar is the collect-from-source half of that pairing.
What to check before you sign
Get concrete answers on these seven points before committing to any provider. Vague answers here are the most common cause of a dataset that looks right in a sample and disappoints in production.
Deduplication method. Ask how the same role posted to five boards resolves to one record, and what happens when the company name or location is written differently on each.
Normalization standards. Ask which taxonomy the provider maps to. The O*NET-SOC 2019 taxonomy covers 1,016 occupational titles built on 867 detailed SOC occupations and encompasses more than 55,000 job titles, and several providers here map to it.
Archive depth and availability. Confirm the archive start date and whether history is included in your license or priced separately.
Revisit and closure logic. Ask how often a posting is rechecked and what evidence marks it closed. This is the single biggest driver of ghost-listing noise in a dataset.
Salary field coverage. Salary disclosure is uneven and moving. The EU Pay Transparency Directive required member states to transpose by June 7, 2026, and employment law firm Jackson Lewis has tracked the resulting employer obligations, including stating a pay range in the ad or before interview. Indeed Hiring Lab data reported in September 2026 showed Italy at 61 percent of postings including pay after rising from 26 percent a year earlier, against 60 percent in the UK, 43 percent in France, 18 percent in Spain, and 14 percent in Germany. Ask what share of records in your target geography actually carry a salary.
Sample terms. Ask for a sample that matches your real query, not a curated slice. Providers offering self-service samples or free tiers make this easy to test.
Pricing behavior at scale. Per-record pricing that looks cheap at 100,000 records can behave very differently at 50 million. Ask for the cost at the volume you expect in year two.
On compliance, two US decisions do most of the framing. In hiQ Labs v. LinkedIn, the Ninth Circuit held in 2019, and reaffirmed in April 2022, that scraping publicly available data likely does not constitute access without authorization under the Computer Fraud and Abuse Act. In Van Buren v. United States, the Supreme Court held in June 2021 that exceeding authorized access means entering off-limits areas of a system rather than using permitted data for an improper purpose. The practical takeaway is narrower than headlines suggest. The CFAA is not the main constraint on collecting public data, but a site's terms of service can still be enforced through contract law. This is context, not legal advice, and your counsel should review any collection program.

Frequently asked questions
What is job posting data used for?
Job posting data powers labor market research, workforce planning, compensation benchmarking, competitive hiring intelligence, sales prospecting, and product features inside HR technology platforms. Central banks and academics use it for real-time labor analysis, and it is also collected as training data for AI models that need structured employment text at scale.
How much does job posting data cost?
Published pricing ranges from free developer tiers to six-figure annual licenses. Techmap prices at roughly $1 per 1,000 postings with historical country datasets from around $2,400, TheirStack starts at about $59 a month above its free tier, and Xverum's marketplace listing shows pricing from $2.50 per 1,000 records up to $100,000 a year. Enterprise analytics providers like Lightcast, LinkUp, and Revelio Labs quote on request. Ficstar's custom collection is the premium managed option, priced on the scope of the work rather than by record count. Our guide to web scraping cost breaks down what actually drives the number.
What is the difference between a job posting dataset and a job posting API?
A dataset is a bulk delivery of records, usually as files on a schedule, suited to analysis and historical work. An API returns records on request, suited to applications that need specific queries in real time. Most providers here sell both, and the choice usually comes down to whether you are populating a warehouse or powering a live feature.
How do providers deduplicate job postings?
Most match on a normalized combination of job title, company, and location within a time window, sometimes with entity resolution across company identifiers. Lightcast, for example, documents matching normalized title, company, and location across a 60-day window. Ask any provider which fields the match key uses, because a match on title and company alone will collapse genuinely distinct roles in different cities.
Can you get job posting data from Indeed and LinkedIn?
Yes, though both actively restrict automated collection, and coverage of them differs sharply between providers. Some providers deliberately avoid aggregators and source only from employer career pages. At Ficstar, we maintain reliable access to sites that block or throttle collection, including Indeed and LinkedIn Jobs, and our crawlers page past the result caps some boards impose.
Which job posting data provider is best for labor market research?
Lightcast, LinkUp, and Revelio Labs are the strongest fits. Lightcast has the deepest US history at January 2010 and maps to O*NET-SOC occupations. LinkUp sources from employer career sites with records back to 2007. Revelio Labs publishes labor statistics benchmarked against BLS releases, which makes its methodology easier to interrogate than most.
Which job posting data provider is best for sales intelligence?
PredictLeads, TheirStack, and Signalbase are built for this use case. PredictLeads ties openings to company records across 2.7 million companies, TheirStack pairs postings with technographics and publishes same-day freshness rates, and Signalbase is designed for outreach timing rather than corpus analysis.
Should I buy a job postings dataset or collect the data myself?
License a dataset when a vendor already covers your sources and fields, you need deep history immediately, or you want enrichment layers you would not build. Collect from the source when you need specific sources, fields, cadence, or format that no packaged product offers, when you need access to sites that block scrapers, or when the data needs to be exclusively yours. A fully managed collector like Ficstar covers that path without you running crawlers in-house.
Get job posting data collected from your sources
If you have looked at packaged datasets and none of them line up with the sources, fields, or cadence you actually need, we can show you what direct collection looks like on your requirement rather than on a demo. Our free trial is real data collection, typically two weeks and longer for complex projects, where we collect from the sources you specify and deliver the records in your format so you can evaluate quality before committing to anything.
Start Your Free Trial and tell us which sources and fields your project depends on.



Comments