Dark grid background with the word PROTOCOL in white monospace type inside a dashed border, positioned below center

How to evaluate a job posting API: a reproducible test protocol for coverage, freshness, history, and licensing

September 8, 2026/Charles·16 min read

How should you evaluate a job posting API?

The simplest way to evaluate a job posting API is to test it against something you can see for yourself. Pick a set of employers whose public career pages you can read directly, then compare what the API shows against what those pages show.

That comparison will set the precedent for every other evaluation that follows, such as what the vendor counts, how quickly new postings appear, whether edits and removals propagate, what survives after a job closes, and what you're actually licensed to do with it.

Evaluation failures often trace back to the buyer accepting a headline number without asking what it measured. For example: A job count can mean raw listings, active postings, deduplicated opportunities, or all-time observed records, and each definition produces a different total for the same employer. "Updated daily" can describe an internal crawl schedule rather than a measured propagation lag. "Historical data" can mean a broad date filter on publication timestamps rather than retained past states.

These are not edge cases. Headline counts can be misleading when the measurement is not defined. This article breaks the evaluation into nine questions, each with a concrete test. It starts with what coverage means when the denominator is undefined, moves through freshness measurement, historical data semantics, duplicate handling, field inspection, and licensing, and ends with a reproducible protocol you can run yourself against any provider's trial.

What does "coverage" mean for job posting data?

A vendor may report a single coverage number without defining what it counts. Before comparing APIs, separate these four measurements.

Raw posting count is the number of records returned before any deduplication. It depends on the query, the timestamp, the filters applied, the pagination method, and the source scope. Two vendors returning different totals for the same employer may simply be counting different things.

Active posting count is the subset that meets an explicitly stated "currently public and open" rule at the time of measurement. A count without that rule is not actionable. Expired postings reachable by direct link, jobs in draft status, and postings removed from the board but not yet flagged as closed can all inflate or deflate the number depending on how the vendor defines "active."

Deduplicated opportunity count groups raw records under a stated identity rule. Greenhouse, for example, distinguishes a unique job-post ID from an underlying internal_job_id for the hiring requisition. A single requisition can generate multiple postings. Counting postings and counting opportunities are different measurements.

All-time observed count is every unique record seen during a specified tracking period or retained in a historical corpus. Without a start date, an end date, a retention policy, and a record identity rule, "all-time" is undefined.

Coverage also varies across employers, geographies, roles, and sources. A vendor that finds every posting at one large employer and none at nine small ones looks excellent in aggregate and poor under an employer-weighted measure. The protocol should publish both: micro source-board recall across all eligible postings and a macro average of employer-level recall.

Source-board recall measures the share of postings on a specific ATS board that the candidate API can return. This is deliberately narrower than "total company coverage" or "total market coverage." A Greenhouse board is an excellent reference for the postings that an employer exposes through Greenhouse, but it's not proof that the board represents every hiring channel the employer uses.

For teams already working through a general provider evaluation framework, the distinction above is where job-posting APIs diverge from person or company enrichment. The denominator is harder to define, and a vendor that doesn't define it is asking the buyer to accept a number without a unit.

How do you measure job posting API freshness?

When a vendor says their data is fresh, that could mean three different things: how fast they find new postings how fast they pick up changes to existing ones how fast they stop showing jobs that have been taken down

Discovery lag is the delay between a posting first becoming observable on a public source board and the candidate API first returning a record for it. A vendor that says "updated daily" is describing an internal process, not a measured lag. The only way to establish discovery lag is to observe both endpoints on the same clock.

Refresh lag is the delay before source-side edits propagate. A posting's title, location, compensation, or description can change after publication. If the candidate API is tested only for new-post discovery, refresh behavior remains unknown.

Removal lag is the delay between a posting ceasing to be publicly listed on the source and the candidate API ceasing to expose it as active. A posting can temporarily disappear from a source due to ATS maintenance or rate limiting. The protocol section below specifies how to confirm a genuine removal versus a transient disappearance.

The timestamp semantics of the sources themselves are part of the problem. Greenhouse documents first_published and updated_at fields on individual jobs. Ashby defines publishedAt as the time a job was last published, meaning a republished job resets the timestamp. Google's structured data specification defines datePosted as the original employer posting date. Treating those timestamps as interchangeable would produce a flawed freshness test.

A stated refresh cadence is not a measured lag. And a measured lag without a sample size, observation window, and percentile distribution is not a reproducible result.

What counts as historical job posting data?

Vendors often describe date filters as historical data. They're not the same thing. A parameter like "any time" narrows a result set by publication date. It doesn't prove that the vendor observed the posting when it was published, that the record survived closure, or that earlier versions are recoverable.

Each type proves something different:

Historical representationWhat it actually proves
Current snapshotWhat the API returned at one observation time. Proves nothing about earlier states unless snapshots were retained.
Snapshot archiveRepeated states captured over time. Useful for reconstructing change, but resolution is bounded by capture cadence.
Event historyExplicit create, update, unpublish, and reopen events with documented time semantics.
Retrospective backfillOld source records loaded after the fact. Potentially valuable, but does not prove contemporaneous observation.
Auditable historical recordA retained past state with clear source/event time and observation/ingestion semantics. The highest bar when longitudinal analysis matters.

Some source systems make the distinction visible. Ashby's public endpoint explicitly describes itself as returning currently published postings. Lever exposes published postings and hides non-published states from its public jobs API. USAJOBS separately documents Historic JOAs as a distinct surface from its current search. That kind of semantic separation is what buyers should look for rather than inferring history from a date filter.

The test for historical depth is direct. Take jobs that close during the trial, plus a pre-trial sample of known-closed postings. Ask whether the API can return them and exactly which states and timestamps survive. A vendor that retains history will have an answer. A vendor whose "history" is a broad publication-date filter will not.

How should you test duplicates, reposts, and multi-location jobs?

Job-posting data has four record-identity cases, and each one inflates counts differently.

Exact duplicates are identical records returned more than once, usually from pagination errors or overlapping queries.

Same-source reposts are new posting IDs for an opportunity the employer has reposted. Greenhouse's separate id and internal_job_id give evaluators a source-native way to distinguish a post from the underlying job. Ashby's publishedAt resets on republication, so a new publication timestamp can itself be a repost signal. A new posting ID for a familiar title cannot automatically be treated as either the same event or a new opening without additional evidence.

Multi-location expansion is a single posting associated with multiple locations. Lever represents this as a primary posting location plus an allLocations array. Whether that becomes one record or several in the candidate API changes the count. The protocol should preserve the source's location structure and report both the listing count and the canonical-opportunity count.

Cross-source duplicates are the same opportunity observed on more than one source. Matching should use deterministic evidence first: exact source URL, then source posting ID, then underlying requisition ID where exposed, then employer plus normalized title plus location plus time window. Any heuristic match should carry a confidence flag and manual-review status. Do not silently convert fuzzy pairs into records of truth.

A vendor that advertises multi-source scale but cannot explain whether identical syndicated postings count once or repeatedly has not defined its unit. For a deeper treatment of cross-source reconciliation, the case for multi-source public web data covers the general problem.

Which fields, normalization, and provenance should you inspect?

A field can be present on most records and still be wrong on many of them. Completeness (is the field filled in?) and correctness (does it match the source?) are separate measurements.

Completeness is the share of records where a field is non-null. Report it separately by field and by stratum. A missing value is not automatically an error. A "95% complete" claim that does not list which fields were included or distinguish optional from required fields is not usable.

Correctness (or source agreement) compares returned values against the public reference source. For source-observable fields, compare both normalized and raw values. Score agreement separately from completeness.

Normalization covers job type, workplace mode, location, employer identity, role taxonomy, and compensation units. The test is whether the evaluator can access both the raw source value and the normalized output. An irreversible normalization with no source value and no way to distinguish observed from inferred fields is a provenance gap.

Provenance is the ability to explain where and when a record or value came from. Audit a random sample for source URL or posting ID, collection timestamp, and observation timestamp. A flat normalized record with no lineage cannot be audited.

Schema drift (keys, types, required fields, and enum values changing between releases) can occur in any API. Save a representative payload on every test day and diff it against the previous day. A field present in every response on day one and missing from half of them on day fifteen is a signal worth investigating, whether it reflects a schema change or a sampling difference.

For a pre-integration inspection checklist that applies to enrichment APIs generally, Part 3 of the evaluation series covers schema clarity, integration mechanics, and sourcing. This protocol fills in the job-posting-specific details that the article leaves general.

What licensing rights should a job data buyer verify?

This section is a set of procurement questions to take to the vendor and the buyer's counsel. It's not legal advice.

Technical availability and downstream data rights are separate questions. USAJOBS provides an API containing publicly available job-announcement information, yet its terms expressly restrict renting, leasing, selling, trading, or redistributing API data to third parties as a standalone product or feed. A public API is not an implicit redistribution license.

Before committing to a provider, get written answers to these questions. Vague assurances like "public data" or "licensed data" are not auditable permissions.

CategoryQuestions for the vendor
Scope of grantWhat exact license applies to API response data, separately from the right to call the API? Is internal analytics, CRM enrichment, customer-facing display, redistribution via API or feed, or AI/model training each permitted?
Downstream accessMay subsidiaries, contractors, cloud processors, and end-user customers access the data? Are there field-level or volume limits on downstream exposure?
Retention and terminationWhat may the buyer retain after a posting closes? What must be deleted when the contract ends, and which derived outputs survive?
Derived dataMay the buyer create and retain normalizations, classifications, aggregate metrics, embeddings, or model features?
Source restrictionsWhat source-specific restrictions flow through? Which sources carry attribution requirements or are prohibited from republication? Are source URLs or apply links required on displayed records?
Takedowns and changesWhat happens if a source demands removal? May the vendor change source composition or restrictions mid-term, and what notice applies?
Warranties and indemnityDoes the vendor warrant it has the necessary rights? What indemnity exclusions, caps, and procedures apply?
Provenance and privacyWhat record-level lineage can the vendor provide? Which fields contain personal information, and what DPA or deletion process applies?
Controlling documentWhich instrument governs if the order form, public terms, API documentation, and source-specific terms conflict?

The buyer's goal is to convert each answer into an auditable permission set for the precise downstream use case.

How do you run a fair job posting API trial?

Before making the first API call, freeze your test sample. Use the same employer, geography, role, and source-board categories from the coverage section.

Freeze the sample. Select a stratified set of employers, geographies, roles, and source boards before the trial begins. The vendor should not choose the evaluation employers. Pre-stratify by employer size and geography so a single high-volume employer cannot dominate the results.

Establish the reference surface. For employers with public ATS boards, the board is the denominator for source-board recall. SmartRecruiters, for example, documents an endpoint for active postings published by a company and exposes a releasedAfter time filter. That kind of surface gives the evaluator a countable, timestamped reference.

Run transport and pagination tests. Record HTTP status codes, schema validity, p50/p95/p99 latency, retries, and eventual success rate over at least seven days of representative scheduled calls. Complete repeated full paginations and count duplicate IDs across pages.

Establish pass/fail criteria before the results arrive. Define acceptable thresholds for discovery lag, employer recall, field completeness, and transport reliability before the trial starts, not after. Retain raw test evidence so the evaluation is auditable.

A posting count is a measurement with a specific unit, coverage scope, and observation time. It's not a proxy for total hiring demand or market completeness. When interpreting job-posting volumes, the same limits that apply to workforce analytics measurements apply here: the number describes what the instrument observed, not what exists.

What does a reproducible freshness test look like?

This section presents a public ATS-board monitor that buyers can reproduce. It assumes the vocabulary from the freshness section above and does not redefine it.

Methodology

  • Documentation checked: September 4, 2026
  • Protocol version: 1.0
  • ATS documentation reviewed: Greenhouse, Lever, Ashby, SmartRecruiters
  • Observation window: Proposed protocol; no longitudinal measurements collected
  • Revision log: v1.0, initial publication

ATS reference endpoints

The protocol uses public ATS job-board APIs as the source of truth. The table below lists each platform's public endpoint documentation and relevant terms. The specific employer boards that form the sample are selected in step 1.

ATS platformPublic endpoint documentationPolling and terms notes
GreenhouseJob Board API: returns published jobs with id, internal_job_id, first_published, updated_atNo authentication required for public boards. No published rate limit; respect Retry-After headers.
LeverPostings API: returns published postings with allLocations arrayPublic, read-only, no API key required. No documented rate limit on the public postings endpoint.
AshbyPublic Job Posting API: returns currently published postings; publishedAt resets on republicationPublic, no API key required. Ashby's developer terms apply; review before automated polling.
SmartRecruitersPostings API: returns active postings with releasedAfter filterPublic company postings endpoint requires no authentication. API terms of service govern automated access.

For reference, USAJOBS provides a public API for federal job announcements but its terms expressly restrict standalone redistribution. It is useful as an example of a public API that is not an implicit redistribution license, as discussed in the licensing section above.

The protocol

For postings observable on a fixed set of employer-controlled public ATS boards, how long does the candidate API take to discover, reflect changes to, and stop exposing each posting?

  1. Select a heterogeneous sample of employer ATS boards across Greenhouse, Lever, Ashby, and SmartRecruiters covering several countries, employer sizes, job volumes, and multi-location patterns. Revalidate board membership at probe start because employers can migrate ATS platforms.

  2. Establish a source baseline before beginning candidate-API comparisons. Record every currently listed posting with its source ID, underlying job or requisition ID where exposed, URL, title, raw location structure, publication timestamp, and payload hash.

  3. Poll the source boards and the candidate API on the same UTC clock. A conservative default is one poll per hour per board. If the vendor's claimed freshness is materially below one hour, a one-hour probe cannot validate that claim.

  4. Run for at least 30 consecutive days. Six to eight weeks is preferable to collect a useful number of natural removals and edits rather than only new-post discovery.

  5. For each new source record, record t_source_first_observed. For the corresponding candidate-API record, record t_vendor_first_observed. Observed discovery lag is the difference. Label it "observed" because both endpoints are discretely polled: a one-hour cadence cannot support minute-level precision.

  6. For tracked field changes, hash a chosen set of source fields (title, description, location, compensation where available) and timestamp changes. Observed refresh lag is the interval between the source change and the candidate API reflecting it.

  7. For removals, require a stable 404 or 410, an explicit expiration field, or at least two consecutive successful, fully paginated checks where the job is absent before generating a removal event. A single missing observation is not sufficient, and timeouts, rate limits, or incomplete responses do not count. Label source removal as "no longer publicly listed on the monitored source," not "job filled," unless the source provides stronger status information.

  8. Exclude left-censored observations: a job already present when monitoring begins cannot enter discovery-lag statistics. Report unresolved cases separately, including how long they remained unresolved at trial end. Do not invent completion times. An API that never discovers or removes a posting during the trial must appear in the results, not vanish from them.

  9. Match source records to candidate-API records using deterministic evidence first: exact source job URL, then source posting ID, then underlying requisition ID, then employer plus normalized title plus location plus time window. Flag heuristic matches with a confidence level. Do not silently convert fuzzy pairs into ground truth.

  10. Report distributions, not averages: discovery-lag median and p90, refresh-lag median and p90 when enough edits occur, removal-lag median and p90, source-board recall by employer and ATS, unmatched-case count, duplicate-group characteristics, and the number of censored events. Any metric based on very few observed events should disclose its sample size.

The research that produced this protocol established source feasibility and a test design across the four ATS platforms listed above. It verified that each exposes documented public job-posting surfaces with materially different field and timestamp semantics. No candidate commercial job-data API was queried, and no vendor freshness results have been produced. The observation window above is unfilled because the longitudinal study has not yet been run.

No named-vendor performance table appears in this article. The reusable value is the protocol itself.

data-engineersrevopsvendor-evaluationdata-freshnessjob-dataapi-integration