{
  "version": "https://jsonfeed.org/version/1.1",
  "title": "Hiring Index",
  "home_page_url": "https://hiringindex.org/blog",
  "feed_url": "https://hiringindex.org/feed.json",
  "description": "Notes on the job market, and on turning job postings into data you can trust.",
  "language": "en",
  "items": [
    {
      "id": "https://hiringindex.org/blog/how-we-measure-posting-age",
      "url": "https://hiringindex.org/blog/how-we-measure-posting-age",
      "title": "How we measure posting age, and what it does not mean",
      "summary": "The definition is narrow on purpose. A wider one would be more impressive and less true.",
      "content_html": "<p>Posting age is the number that runs through everything we publish. It deserves an exact definition, because the loose version of it would let us claim something we cannot support.</p>\n<h2>The definition</h2>\n<p><strong>Age is the number of days between the date the employer published the listing and today.</strong> Nothing else.</p>\n<h2>What it does not mean</h2>\n<ul><li>It does not mean we have watched the listing continuously since that date.</li><li>It does not mean the role is still open. It means the listing was still visible when we last checked.</li><li>It does not mean nobody is reading applications. Intent is not observable from outside.</li></ul>\n<p><em>Some vendors describe a similar number as \"days live\", which implies continuous observation. If we said that, we would be claiming a history we may not have for every source. Age is the honest word.</em></p>\n<h2>Where the date comes from</h2>\n<ol><li>A structured publication date from the source, where one exists. This is the best case and the most common.</li><li>A date parsed from the listing text where the source states one plainly.</li><li>Where a source gives only a coarse indication such as \"30+ days ago\", we mark the field approximate rather than inventing a day.</li></ol>\n<p>Where none of these is available, the field is absent. It is not estimated, and it is not filled with the date we first saw the listing dressed up as a publication date.</p>\n<h2>Why sixty days</h2>\n<p>The site marks listings older than sixty days. The threshold is a convention, not a discovery. It sits above the typical time-to-fill for most roles and below the point where a listing is obviously abandoned. Your market may differ; the raw number is always shown next to the marker so you can apply your own line.</p>\n<h2>The known biases</h2>\n<ul><li><strong>Republishing.</strong> An employer who deletes and re-posts resets the age. Their listings look fresher than they are, and we cannot always tell.</li><li><strong>Source mix.</strong> Applicant tracking systems remove closed roles promptly; some boards do not. Age distributions differ by source, not only by market.</li><li><strong>Missing dates.</strong> Cohorts where many postings lack a date have age computed from the subset that has one, which may not be representative.</li></ul>\n<h2>What we would need to say more</h2>\n<p>To claim a listing has been continuously open, we would need a complete observation history for every source, with gaps accounted for. We publish what we can measure and name the gap rather than closing it with an assumption.</p>",
      "date_published": "2026-08-28T09:00:00.000Z",
      "tags": [
        "method",
        "posting age",
        "transparency"
      ]
    },
    {
      "id": "https://hiringindex.org/blog/what-a-job-data-api-actually-gives-you",
      "url": "https://hiringindex.org/blog/what-a-job-data-api-actually-gives-you",
      "title": "What a job data API actually gives you",
      "summary": "The pitch is “millions of listings”. The work is in the fields nobody advertises.",
      "content_html": "<p>Every provider in this market leads with a count: millions of postings, hundreds of thousands of companies, dozens of sources. The count is the least useful thing about the product, because raw postings are cheap and the hard parts are elsewhere.</p>\n<h2>What you are actually buying</h2>\n<ol><li><strong>Collection.</strong> Someone runs the crawlers, handles the blocks, and keeps up when a source changes its markup. This is continuous work, not a one-off.</li><li><strong>Deduplication.</strong> One role cross-posted to five boards must become one record. Without this your counts are inflated by a factor you cannot estimate.</li><li><strong>Normalisation.</strong> Free-text locations into city, region, country. Salary strings into a number, a currency and a period. Titles into something groupable.</li><li><strong>Expiry.</strong> Knowing when a listing came down is harder than knowing when it went up, and more valuable.</li><li><strong>A stable schema.</strong> Fields that mean the same thing next quarter.</li></ol>\n<h2>The fields that separate providers</h2>\n\n<h2>Questions worth asking before you buy</h2>\n<ul><li>What share of postings have a normalised salary, not just a salary field?</li><li>How do you decide two postings are the same role, and what is your error rate?</li><li>Do you keep first-seen timestamps, or only the current snapshot?</li><li>When a listing disappears from the source, how quickly do I learn?</li><li>Where does the data come from — the employer's own systems, or another vendor's feed?</li></ul>\n<p>That last one matters commercially and legally. A provider re-selling another vendor's feed inherits that vendor's risk. Proxycurl, a well-known provider of adjacent data, shut down in July 2025 following litigation from LinkedIn. Where the data comes from is not a detail.</p>\n<h2>What &quot;real time&quot; usually means</h2>\n<p>Rarely what it sounds like. Ask for the distribution, not the claim: what share of postings are discovered within an hour, within a day, within a week. A provider who cannot answer has not measured it.</p>",
      "date_published": "2026-08-26T09:00:00.000Z",
      "tags": [
        "api",
        "job data",
        "buying"
      ]
    },
    {
      "id": "https://hiringindex.org/blog/what-posting-age-tells-you-about-a-market",
      "url": "https://hiringindex.org/blog/what-posting-age-tells-you-about-a-market",
      "title": "What posting age tells you about a market",
      "summary": "Two markets can have identical vacancy counts and behave completely differently. Age is what separates them.",
      "content_html": "<p>Vacancy counts are the standard measure of labour demand and they hide the most useful thing about a market: how fast it moves.</p>\n<h2>Same count, opposite markets</h2>\n<p>Two cities each show 400 open data engineering roles. In the first, the median posting is nine days old and a fifth of listings were published this week. In the second, the median is seventy-one days and almost nothing is new.</p>\n<p>The first market is churning: roles open, fill, and are replaced. The second is silted up — the same roles have been advertised since spring. A candidate treating those as equivalent will waste months, and so will anyone forecasting from the counts.</p>\n<h2>How to read the number</h2>\n\n<p><em>These bands are conventions from looking at a lot of cohorts, not thresholds derived from a study. Use them to ask questions, not to draw conclusions.</em></p>\n<h2>Age and momentum together</h2>\n<p>Age alone can mislead. Pair it with the thirty-day change in openings:</p>\n<ul><li><strong>Young listings, rising count</strong> — genuine expansion. The best signal available.</li><li><strong>Young listings, falling count</strong> — roles are being filled faster than opened. Tightening, not weakening.</li><li><strong>Old listings, rising count</strong> — postings are accumulating without resolving. Often the first sign of a freeze that has not been announced.</li><li><strong>Old listings, falling count</strong> — retreat. Roles are being withdrawn rather than filled.</li></ul>\n<p>The third combination is the one worth watching, because it looks healthy on a vacancy count and is not.</p>\n<h2>Why almost nobody publishes it</h2>\n<p>Computing age needs the publication date on every listing and a consistent definition of what counts as one posting. Providers who assemble their index from boards inherit ingestion dates rather than publication dates, which makes everything look fresh. It is not that the number is hard to understand — it is that it is easy to compute wrongly and hard to compute honestly.</p>",
      "date_published": "2026-08-25T09:00:00.000Z",
      "tags": [
        "posting age",
        "labour market",
        "analysis"
      ]
    },
    {
      "id": "https://hiringindex.org/blog/building-a-job-board-what-you-actually-need",
      "url": "https://hiringindex.org/blog/building-a-job-board-what-you-actually-need",
      "title": "Building a job board: what you actually need",
      "summary": "The listings are the easy part. Everything that makes a board worth visiting is downstream of them.",
      "content_html": "<p>Job boards look like a solved problem: get listings, show listings. Every operator discovers in month two that the listings were never the problem.</p>\n<h2>The five things that decide whether it works</h2>\n<ol><li><strong>Expiry.</strong> Nothing kills a niche board faster than dead listings. Users forgive a small index; they do not forgive applying to something that closed in March.</li><li><strong>Deduplication.</strong> Four copies of one role makes a board of 400 feel like a board of 100 and look careless.</li><li><strong>A defensible slice.</strong> \"Jobs\" is not a market. \"Rust jobs in Europe with visa sponsorship\" is.</li><li><strong>Search that matches how people think.</strong> People search by outcome — remote, salary, visa — not by taxonomy.</li><li><strong>A reason to come back.</strong> Alerts, saved searches, a weekly digest. Without one, every visit is a new acquisition.</li></ol>\n<h2>The economics nobody mentions</h2>\n<p>The default assumption is that a board sells listings to employers. In practice, small boards usually start on the other side: aggregators supply listings free and pay the board per click. That inverts the model — the board is a traffic business, not a marketplace — and it changes what you should optimise.</p>\n<p><em>It also means \"buy job data\" and \"make money from job listings\" are separate decisions. Plenty of boards never pay for data at all.</em></p>\n<h2>A build order that works</h2>\n<ol><li>Pick a slice narrow enough that you can name every employer in it.</li><li>Get listings into a database, however crudely. Manual is fine at first.</li><li>Solve expiry before design. A board of fifty live roles beats a board of five hundred where a third are dead.</li><li>Add one retention mechanism — an email alert is enough.</li><li>Only then worry about how it looks.</li></ol>\n<h2>What to buy and what to build</h2>\n\n<h2>The mistake almost everyone makes</h2>\n<p>Launching with the largest index you can assemble. A big index with stale entries is worse than a small curated one, because trust is the only thing a niche board has, and one dead listing spends more of it than ten live ones earn.</p>",
      "date_published": "2026-08-23T09:00:00.000Z",
      "tags": [
        "job boards",
        "building",
        "product"
      ]
    },
    {
      "id": "https://hiringindex.org/blog/how-to-read-a-salary-range",
      "url": "https://hiringindex.org/blog/how-to-read-a-salary-range",
      "title": "How to read a salary range without fooling yourself",
      "summary": "The number in a job ad is a negotiating position, a legal artefact, and a filter — usually all three at once.",
      "content_html": "<p>A range in a job ad is not a measurement. It is a decision someone made, under constraints you cannot see. Reading it well means knowing which constraints were in play.</p>\n<h2>What a range is actually doing</h2>\n<ol><li><strong>Filtering.</strong> The bottom exists to stop applications from people who will decline. It is often below what the employer expects to pay.</li><li><strong>Complying.</strong> Where disclosure is required by law, the range must be posted in good faith — but \"good faith\" is wide, and wide ranges are the common response.</li><li><strong>Anchoring.</strong> The top exists to attract, and is frequently reserved for a candidate who does not exist.</li></ol>\n<h2>The width tells you more than the midpoint</h2>\n<p>A range of $120,000–$135,000 is a real band for one level of one role. A range of $90,000–$190,000 is not a band; it is several levels advertised as one requisition, or a compliance minimum met without commitment.</p>\n<p>Practical rule: if the top is more than about 40% above the bottom, treat the posting as covering multiple levels. Ask which level you are being considered for before discussing numbers.</p>\n<h2>The disclosure rate is context you need</h2>\n<p>When you see a median for a role and city, ask what share of postings disclosed anything. A median computed from 12% of a cohort is a different object from one computed from 70%, because the postings that disclose are not a random sample — they skew towards regulated markets, larger employers, and roles where the employer feels competitive.</p>\n<p><em>This is why every median on this site is printed next to the share of postings that disclosed a salary. A number without that context invites a false conclusion.</em></p>\n<h2>What ranges systematically omit</h2>\n<ul><li><strong>Equity.</strong> Where it exists, it is often the largest variable and almost never in the ad.</li><li><strong>Bonus.</strong> Sometimes included in the top of the range, sometimes not, rarely stated which.</li><li><strong>Level.</strong> The same title spans very different bands between employers.</li><li><strong>Location adjustment.</strong> For remote roles, a single range may be adjusted downward after an offer, based on where you live.</li></ul>\n<h2>Comparing across countries</h2>\n<p>Do not. A median in Berlin and a median in Austin are separated by currency, tax treatment, employer social contributions, healthcare, holiday entitlement and pension. Comparing the raw numbers produces a conclusion that is confidently wrong.</p>\n<p>Compare within a country, and preferably within a metro. That is why our city comparisons are restricted to pairs inside the same country — a comparison across borders would look informative and mislead.</p>\n<h2>The single most useful question</h2>\n<p>“What is the band for this level, and where in it would you expect someone with my experience to land?” It converts a filter into a measurement, and it is answerable by any employer acting in good faith.</p>",
      "date_published": "2026-08-21T09:00:00.000Z",
      "tags": [
        "salary",
        "negotiation",
        "pay transparency"
      ]
    },
    {
      "id": "https://hiringindex.org/blog/deduplicating-job-postings",
      "url": "https://hiringindex.org/blog/deduplicating-job-postings",
      "title": "Deduplicating job postings is harder than it looks",
      "summary": "Exact matching finds a fraction. Fuzzy matching merges roles that are genuinely different. Here is the shape of the problem.",
      "content_html": "<p>Every team that touches job data writes a deduplicator, and every one of them underestimates it. The naive version takes an afternoon and is wrong in ways you only discover months later, in aggregate numbers you have already reported.</p>\n<h2>Why exact matching fails</h2>\n<p>The same role reaches different boards through different pipelines. Titles get suffixes: <code>Senior Engineer</code>, <code>Senior Engineer (Remote)</code>, <code>Senior Engineer - Berlin</code>. Descriptions get truncated at different lengths. HTML gets stripped differently. Salary appears in one copy and not another.</p>\n<p>Hashing the description finds the small share of copies that travelled without modification. Everything else slips through.</p>\n<h2>Why fuzzy matching over-merges</h2>\n<p>Loosen the criteria and you merge things that should stay apart:</p>\n<ul><li>A company hiring three Backend Engineers in one city — three requisitions, near-identical text.</li><li>The same title at the same company in different offices, which is genuinely different work.</li><li>A role reposted after six months, which is a new requisition, not a duplicate.</li><li>Agency and employer versions of one role, where merging loses the fact that an agency is involved.</li></ul>\n<h2>What actually works</h2>\n<p>A composite key, applied in order, with the strongest evidence first:</p>\n<ol><li><strong>Source identity.</strong> If two records share an applicant tracking system id, they are the same posting. Free and certain when available.</li><li><strong>Apply URL.</strong> Normalise it — strip tracking parameters — and identical destinations are the same role.</li><li><strong>Employer + normalised title + normalised location.</strong> Strong, but needs a tie-break for genuine multi-seat requisitions.</li><li><strong>Description similarity.</strong> Shingled hashing over the description, as confirmation rather than as primary evidence.</li><li><strong>Time window.</strong> Two matching records six months apart are two requisitions, not one.</li></ol>\n<h2>The multi-seat problem</h2>\n<p>There is no clean answer to a company advertising four identical seats. Merging understates demand; keeping them separate overstates it if they are actually one posting duplicated by a board. The honest approach is to pick a rule, publish it, and stay consistent — so at least your series is comparable with itself.</p>\n<p>Ours: identical employer, title and location inside a short window collapse to one record, with every source retained. It undercounts multi-seat hiring. We would rather undercount predictably than overcount invisibly.</p>\n<h2>How to know yours is wrong</h2>\n<p>Take one large employer you can verify by hand. Count what your pipeline says they have open. Count what their careers page says. If those differ by more than a few percent, the difference is your deduplication, and it is applying to every number you produce.</p>",
      "date_published": "2026-08-19T09:00:00.000Z",
      "tags": [
        "deduplication",
        "engineering",
        "job data"
      ]
    },
    {
      "id": "https://hiringindex.org/blog/structured-data-for-job-postings",
      "url": "https://hiringindex.org/blog/structured-data-for-job-postings",
      "title": "Structured data for job postings, and why most implementations are wrong",
      "summary": "JobPosting markup is how a listing gets into Google Jobs. The common errors are all in the same three fields.",
      "content_html": "<p>If you publish job listings and want them surfaced in job search results, you need <code>JobPosting</code> structured data. The specification is short. Most implementations still get the same handful of things wrong.</p>\n<h2>The minimum that actually works</h2>\n\n<h2>The three fields people get wrong</h2>\n<ol><li><strong>datePosted.</strong> Must be when the employer published the role, not when you ingested it. Boards that stamp ingestion time make every listing look fresh, which is exactly the behaviour the field exists to expose.</li><li><strong>Remote roles.</strong> A remote job needs <code>jobLocationType: \"TELECOMMUTE\"</code>, and if there is a geographic restriction, <code>applicantLocationRequirements</code>. Putting the company headquarters in <code>jobLocation</code> for a remote role is both wrong and common.</li><li><strong>baseSalary.</strong> If you do not have a salary, omit the field. Emitting a zero, a placeholder, or a made-up range is worse than silence.</li></ol>\n<h2>validThrough is not optional in practice</h2>\n<p>Listings without an expiry stay eligible until something removes them. Set <code>validThrough</code>, and remove the markup when the role closes. A listing that outlives the requisition is the single most common complaint about job search results, and it is a self-inflicted one.</p>\n<h2>Things that look helpful and are not</h2>\n<ul><li>Marking up a search results page as a JobPosting. It is a list, not a posting.</li><li>Duplicating markup for the same role on several of your own URLs. Pick a canonical.</li><li>Inflating <code>description</code> with keywords. It is rendered to users.</li></ul>\n<h2>How to check</h2>\n<p>Validate the JSON, then look at the rendered page and ask whether the markup describes what a person sees. Markup that disagrees with the visible page is the category of error that gets a site removed from job results rather than merely ranked lower.</p>",
      "date_published": "2026-08-16T09:00:00.000Z",
      "tags": [
        "seo",
        "structured data",
        "job boards"
      ]
    },
    {
      "id": "https://hiringindex.org/blog/how-to-tell-if-a-job-posting-is-real",
      "url": "https://hiringindex.org/blog/how-to-tell-if-a-job-posting-is-real",
      "title": "How to tell if a job posting is real",
      "summary": "Six signals you can check in under a minute, ranked by how much they actually tell you.",
      "content_html": "<p>Most advice on this subject is vibes. Here is what can actually be checked from the outside, ordered by how much each signal is worth.</p>\n<h2>1. How long it has been up</h2>\n<p>This is the strongest signal available to you, and the cheapest to check. In a survey of 1,000 US job seekers published by Enhancv in March 2026, <strong>27.2% named the age of a listing — active for three months or more — as the most reliable indicator that a role does not exist</strong>. It beat every other signal they were offered.</p>\n<p>The logic is simple. A real requisition has a hiring manager waiting on it, and that person gets impatient. A listing that has sat untouched since spring is either filled, frozen, or was never a real opening.</p>\n<p><em>A long-lived listing is a signal, not a verdict. Senior and specialised roles genuinely take months. Treat ninety days as a reason to ask, not as proof.</em></p>\n<h2>2. Whether the same role keeps reappearing</h2>\n<p>A posting that is taken down and republished every few weeks is doing something other than hiring. Sometimes it is a board refreshing stale inventory. Sometimes it is a company keeping a pipeline warm. Either way, the requisition is not moving.</p>\n<p>You can check this by searching the exact job title plus the company name and looking at the dates on the results. If three different boards show three different posting dates for what is obviously the same role, the freshest date is marketing, not fact.</p>\n<h2>3. Whether it came from the employer or an agency</h2>\n<p>Agency listings are not fake, but they are further from the truth. The agency may be working a real requisition, may be fishing for candidates to place elsewhere, or may be advertising a role that closed a month ago because the listing still generates applications.</p>\n<p>Tell-tale signs: no company name, a description written in the third person about \"our client\", and a salary given as a wide band rather than a range.</p>\n<h2>4. Whether the description describes work or describes a person</h2>\n<p>Real requisitions are written by someone who knows what the job involves. They mention the team, the systems, the first project. Fabricated or recycled listings describe an idealised person: driven, passionate, a self-starter who thrives in ambiguity.</p>\n<p>The test: after reading it, can you say what you would do on the first Monday? If not, nobody has thought about the role concretely.</p>\n<h2>5. Whether the salary is there at all</h2>\n<p>In markets without pay transparency law, most postings omit salary, so its absence tells you little. In markets where disclosure is required — Colorado, New York City, California, and increasingly the EU — a missing range is a real signal that something is off about the listing.</p>\n<h2>6. Whether the application route is the company&#39;s own system</h2>\n<p>An application link that goes to the company's own applicant tracking system is a good sign: someone configured a requisition in software the company pays for. A link to a generic form, a personal email address, or a third-party site that then asks for your details again is weaker.</p>\n<h2>What none of this tells you</h2>\n<p>None of these signals reveal intent. A company can have a genuine, funded, urgent requisition and still ghost you. A listing can be four months old because the hiring manager left and nobody cleaned up. The signals tell you where to spend your effort, not who is honest.</p>\n<p>Our own view, which we publish because it constrains what we claim: we measure the date a posting was published and how long it has been visible. We do not know whether anyone is reading the applications, and we do not pretend to.</p>\n<h2>A one-minute check</h2>\n<ol><li>Look at the posting date. Older than sixty days without a re-post? Note it.</li><li>Search the exact title plus company. Multiple dates for one role means the dates are unreliable.</li><li>Read for the first Monday. If you cannot picture it, the role may not be defined.</li><li>Check who is hiring: the employer or an agency.</li><li>Check where the apply button goes.</li><li>If the market requires pay disclosure and there is none, downgrade it.</li></ol>",
      "date_published": "2026-08-14T09:00:00.000Z",
      "tags": [
        "ghost jobs",
        "job search",
        "signals"
      ]
    },
    {
      "id": "https://hiringindex.org/blog/applicant-tracking-systems-landscape",
      "url": "https://hiringindex.org/blog/applicant-tracking-systems-landscape",
      "title": "The applicant tracking system landscape, from the outside",
      "summary": "Most job postings pass through a dozen or so systems. Knowing which changes what you can collect.",
      "content_html": "<p>When a company posts a job, it almost always does it inside an applicant tracking system. The system publishes a careers page, syndicates to boards, and receives applications. For anyone collecting job data, the ATS is the layer that matters, because it is the closest thing to a source of truth.</p>\n<h2>Why the ATS layer is better than boards</h2>\n<ul><li><strong>Timing.</strong> A posting exists there first, before any board ingests it.</li><li><strong>Fidelity.</strong> The description is unedited and untruncated.</li><li><strong>Identity.</strong> Requisition ids make deduplication tractable.</li><li><strong>Expiry.</strong> When the requisition closes, the listing usually disappears immediately, rather than lingering on a board.</li></ul>\n<h2>The practical consequence</h2>\n<p>A dataset assembled from boards and one assembled from careers systems differ systematically, not randomly. The board dataset is later, more duplicated, and lags on expiry. The careers-system dataset is fresher and cleaner but misses employers who post only to boards — which is a real and sizeable group at the smaller end.</p>\n<p>Neither is complete. Anyone claiming complete coverage of a national job market is describing an ambition.</p>\n<h2>What varies between systems</h2>\n\n<h2>What this means for the fields you get</h2>\n<p>The completeness of any field in a job dataset is not a property of the provider — it is a property of the mix of systems underneath. A provider whose coverage skews to systems with structured salary fields will report a higher disclosure rate than one whose coverage skews elsewhere, without either being wrong.</p>\n<p>Which is another reason to ask a provider not just what they have, but where it came from.</p>",
      "date_published": "2026-08-12T09:00:00.000Z",
      "tags": [
        "ats",
        "job data",
        "engineering"
      ]
    },
    {
      "id": "https://hiringindex.org/blog/what-counts-as-one-posting",
      "url": "https://hiringindex.org/blog/what-counts-as-one-posting",
      "title": "What counts as one posting",
      "summary": "Every number on this site depends on this definition, so here it is in full.",
      "content_html": "<p>Counting sounds trivial until you try. The same role appears on several boards; one requisition may advertise four seats; an agency and an employer may both advertise the same job. Each choice changes every downstream number.</p>\n<h2>Our rule</h2>\n<p><strong>One posting is one role at one employer in one location, regardless of how many places it appears.</strong></p>\n<h2>What follows from it</h2>\n\n<h2>The one we get wrong on purpose</h2>\n<p>Multi-seat requisitions are undercounted. A company opening one requisition for six warehouse staff counts once here, and as six openings in official vacancy statistics. We chose predictable undercounting over invisible overcounting, because a duplicate-inflated count cannot be corrected after the fact by anyone reading it.</p>\n<p>If you are comparing our counts to an official series, this is the main reason they will differ, and the gap will be largest in high-volume hiring.</p>\n<h2>How to check us</h2>\n<p>Pick an employer whose careers page you can read. Count their open roles by hand. Compare with what we report. If the numbers diverge by more than a few percent, we want to know — that is exactly the sort of correction the <a href=\"/feedback\">feedback form</a> exists for.</p>",
      "date_published": "2026-08-11T09:00:00.000Z",
      "tags": [
        "method",
        "deduplication",
        "transparency"
      ]
    },
    {
      "id": "https://hiringindex.org/blog/why-job-counts-differ-between-sources",
      "url": "https://hiringindex.org/blog/why-job-counts-differ-between-sources",
      "title": "Why job counts differ between sources, by a lot",
      "summary": "Two providers report the same market and differ threefold. Usually none of them is lying.",
      "content_html": "<p>Ask three providers how many software engineering roles are open in Germany and you will get three answers that are not close. The gap is almost never dishonesty. It is definitions.</p>\n<h2>The five choices that produce the gap</h2>\n<ol><li><strong>Deduplication.</strong> Counting listings rather than roles inflates a market by however much cross-posting happens in it — which varies by country and by seniority.</li><li><strong>Expiry.</strong> A provider that removes a listing the day it disappears reports fewer roles than one that keeps it for thirty days. Both are defensible; only one is comparable to a vacancy statistic.</li><li><strong>Multi-seat requisitions.</strong> One requisition for six seats is one or six depending on the rule.</li><li><strong>Agency listings.</strong> Include them and the count rises; the same underlying role may be advertised by an employer and two agencies.</li><li><strong>Source mix.</strong> A careers-page index and a board index cover different employers. Neither is a superset.</li></ol>\n<h2>What to ask before comparing</h2>\n<ul><li>Are you counting listings or roles?</li><li>When a listing disappears from the source, how long until it leaves your count?</li><li>Do you include agency-sourced postings, and can I exclude them?</li><li>What is your source mix — careers systems, boards, or a purchased feed?</li></ul>\n<p>A provider who cannot answer these has not decided, which means their number drifts as their pipeline changes.</p>\n<h2>The practical advice</h2>\n<p>Do not mix sources in one series. A chart that switches provider halfway will show a step change that looks like a labour market event and is a definitional artefact. Pick one, understand its rules, and treat the absolute level as less trustworthy than the direction.</p>\n<h2>And about official statistics</h2>\n<p>Government vacancy series count differently again — usually seats, from employer surveys, with a lag of weeks. Job posting data is faster and broader; official series are slower and more consistent. They answer different questions and the gap between them is not an error in either.</p>",
      "date_published": "2026-08-09T09:00:00.000Z",
      "tags": [
        "job data",
        "counting",
        "method"
      ]
    },
    {
      "id": "https://hiringindex.org/blog/what-ghost-jobs-are-and-what-the-research-says",
      "url": "https://hiringindex.org/blog/what-ghost-jobs-are-and-what-the-research-says",
      "title": "What ghost jobs are, and what the research actually says",
      "summary": "The number everyone quotes is 20%. Here is where it comes from, what it measures, and why the other estimates disagree.",
      "content_html": "<p>“Ghost job” means a posting advertised without an intention to fill it. Companies do it to build a talent pipeline, to signal growth, to satisfy an internal policy that a role be advertised before an internal candidate takes it, or because nobody remembered to take the listing down.</p>\n<p>The figure you see quoted is usually “one in five”. It has a specific source, and knowing that source tells you what the number does and does not cover.</p>\n<h2>Where the 20% comes from</h2>\n<p>The academic reference is Hunter Ng, <em>Why is it so hard to find a job now? Enter Ghost Jobs</em>, submitted to arXiv in October 2024. Using a dataset from Glassdoor and an LLM-BERT classification approach, the study finds that <strong>up to 21% of job ads may be ghost jobs</strong>, more so in specialised industries and larger firms.</p>\n<p>The paper goes further than the headline: it argues ghost postings help explain the recent disconnect in the Beveridge curve — the long-standing relationship between vacancies and unemployment that stopped behaving as expected over the past fifteen years.</p>\n<blockquote>Up to 21% of job ads may be ghost jobs, and this is particularly prevalent in specialized industries and in larger firms.</blockquote>\n<p>Note the hedge: <em>up to</em>. The popular retelling drops it and reports a flat “1 in 5”.</p>\n<h2>Why other estimates disagree</h2>\n\n<p>These are not competing measurements of the same thing. A survey asking managers whether they have ever posted a role they were not filling will always produce a larger number than a method that classifies individual listings, because one manager admitting the practice covers many postings.</p>\n<h2>What changed, and when</h2>\n<p>The topic peaked as news rather than as a phenomenon. Measured by attention on Hacker News — one reasonable proxy for what technical audiences are arguing about — stories with the phrase in the headline collected 1,484 points across 38 stories in 2024, and 1,403 points across only 16 stories in 2025: fewer articles, far more engagement each. Through August 2026, 22 stories collected 433 points.</p>\n<p>Read that as the topic moving from news to background knowledge. People stopped upvoting explanations because they no longer need one.</p>\n<h2>The part that gets ignored</h2>\n<p>Almost every well-known ghost job story is about the effect on candidates. The most-discussed stories in this space are about something else entirely: fake job offers as an <em>attack vector</em>. The single highest-scoring posting-related story on Hacker News is about a fabricated job offer that led to the Axie Infinity breach and a loss of over $600m. Others cover malware delivered inside take-home assignments, and fabricated interviews treated as securities fraud.</p>\n<p>If you are a security team rather than a job seeker, the ghost job problem looks completely different, and nobody in the job data business is looking at it.</p>\n<h2>What we can and cannot see</h2>\n<p>We publish how long postings stay open, because that is measurable from the outside. We cannot see inside a requisition. Anyone claiming to detect intent from a listing is selling a model of intent, not a measurement — and the honest version of that product would report a probability, not a verdict.</p>",
      "date_published": "2026-08-07T09:00:00.000Z",
      "tags": [
        "ghost jobs",
        "research",
        "labour market"
      ]
    },
    {
      "id": "https://hiringindex.org/blog/why-so-few-job-ads-show-pay",
      "url": "https://hiringindex.org/blog/why-so-few-job-ads-show-pay",
      "title": "Why so few job ads show pay, and where that is changing",
      "summary": "Disclosure is a policy artefact, not a courtesy. The map explains most of the variation.",
      "content_html": "<p>The share of postings that publish a salary varies by a factor of five depending on where the job is and who is hiring. Almost none of that variation is about generosity.</p>\n<h2>The three real reasons</h2>\n<ol><li><strong>Law.</strong> Where disclosure is mandated, rates jump immediately and stay high. Where it is not, most employers omit it.</li><li><strong>Negotiating position.</strong> An employer that does not publish keeps the option of paying different people differently for the same work. Publishing removes that option, which is the point of the laws.</li><li><strong>Internal equity risk.</strong> Publishing a range for a new hire tells existing staff what the company thinks the role is worth. Many employers would rather not have that conversation.</li></ol>\n<h2>Where disclosure is required</h2>\n<p>The direction of travel is one way. Several US states and cities require a range in the posting; the EU pay transparency directive obliges member states to require pay information to candidates before interview, with national implementations landing through 2026. Enforcement varies more than the statutes do.</p>\n<p><em>Requirements change and vary by employer size and role location. Treat any list of jurisdictions as a starting point for checking, not as legal advice.</em></p>\n<h2>What happens to the data when a law lands</h2>\n<p>Two things, and only one is good. Disclosure rates rise sharply — that is the intended effect. Ranges also widen, because a wide band satisfies the requirement while preserving flexibility. The result is more data of lower precision, which is still an improvement on no data.</p>\n<h2>What this means when you read our numbers</h2>\n<p>A cohort in a disclosure jurisdiction has a higher disclosure rate and wider ranges. A cohort without one has a lower rate and narrower ranges from a self-selected group of employers. The medians are not directly comparable, and we print the disclosure rate next to every one of them so the difference is visible rather than hidden.</p>",
      "date_published": "2026-08-05T09:00:00.000Z",
      "tags": [
        "pay transparency",
        "salary",
        "regulation"
      ]
    },
    {
      "id": "https://hiringindex.org/blog/why-the-same-job-appears-on-five-sites",
      "url": "https://hiringindex.org/blog/why-the-same-job-appears-on-five-sites",
      "title": "Why the same job appears on five different sites",
      "summary": "Cross-posting, syndication and scraping produce four copies of one role. Here is how to tell which is the original.",
      "content_html": "<p>You search for a role, get eight results, and slowly realise they are the same job. This is not a glitch. It is how the market is wired.</p>\n<h2>Four ways one posting becomes many</h2>\n<ol><li><strong>The employer cross-posts.</strong> A recruiter publishes to their own careers page and to two or three boards, by hand or through their applicant tracking system.</li><li><strong>Boards syndicate.</strong> Large aggregators buy or exchange feeds. One listing enters the network and appears on partner sites automatically, sometimes with the description rewritten.</li><li><strong>Agencies repost.</strong> A staffing firm working the requisition advertises it under their own name, without the employer.</li><li><strong>Scrapers copy.</strong> Smaller sites populate themselves from bigger ones, so the listing arrives third-hand with whatever errors accumulated on the way.</li></ol>\n<h2>How to find the original</h2>\n<ul><li>Follow the apply link. If it lands on an applicant tracking system under the company's own domain, you are close to the source.</li><li>Compare posting dates. The earliest date is usually the original; later copies inherit the date they were ingested, not the date the role opened.</li><li>Compare descriptions. The original tends to be longer and less tidy. Syndicated copies get truncated and reformatted.</li><li>Look for the requisition number. If one copy has an internal reference code and the others do not, that copy came from the employer's system.</li></ul>\n<h2>Why it matters more than it seems</h2>\n<p>Duplication distorts everything downstream. A job board that counts listings rather than roles will tell you a market is three times larger than it is. Anyone drawing conclusions about hiring demand from raw posting counts is measuring syndication as much as demand.</p>\n<p>It also wastes your time: applying through two copies of one role can look, from the employer's side, like two applications from the same person on the same day.</p>\n<h2>How this gets solved on the data side</h2>\n<p>Deduplication is the least glamorous and most valuable thing a job data provider does. It usually combines a normalised title, the employer identity, the location, and a fingerprint of the description. None of those alone is enough: the same company posts genuinely different roles with identical titles, and the same role gets rewritten between boards.</p>\n<p>We collapse copies into one record and keep every source attached, so you can see where a role appeared and when. The count you see is roles, not listings.</p>",
      "date_published": "2026-07-31T09:00:00.000Z",
      "tags": [
        "duplicates",
        "job boards",
        "job search"
      ]
    },
    {
      "id": "https://hiringindex.org/blog/normalising-job-titles",
      "url": "https://hiringindex.org/blog/normalising-job-titles",
      "title": "Normalising job titles: the problem nobody has solved well",
      "summary": "Titles are marketing, not taxonomy. Every approach to grouping them trades one kind of error for another.",
      "content_html": "<p>Job titles are written to attract candidates and to satisfy internal grading. Neither purpose produces a consistent vocabulary. Any analysis that groups by title is making assumptions, and it is worth knowing which.</p>\n<h2>Four kinds of mess</h2>\n<ol><li><strong>Synonyms.</strong> Software Engineer, Developer, Programmer, SDE, Member of Technical Staff.</li><li><strong>Seniority baked in.</strong> Senior, Staff, Principal, Lead, II, III — inconsistently and not comparably between employers.</li><li><strong>Invented titles.</strong> Growth Ninja, Customer Happiness Hero. Rare but never zero.</li><li><strong>Composite roles.</strong> \"Data Engineer / Analytics Engineer\" is genuinely two roles in one requisition.</li></ol>\n<h2>The three approaches, and what each costs</h2>\n\n<p>There is no fourth option that avoids the trade-off. Anyone who tells you their title normalisation is solved has chosen one of these and stopped mentioning the cost.</p>\n<h2>The seniority trap</h2>\n<p>Stripping seniority to group titles is usually right for demand analysis and usually wrong for salary analysis. \"Engineer\" and \"Staff Engineer\" belong in the same demand bucket and in very different pay buckets. Any system that uses one grouping for both will produce a median that describes nobody.</p>\n<p>We keep seniority as a separate dimension so that a cohort can be grouped one way and split the other.</p>\n<h2>A practical recommendation</h2>\n<p>Use rules for the roles you care about and be honest that everything else is a long tail. A curated list of eighty roles you can defend beats an automatic clustering of eight thousand you cannot explain to a customer who asks why two things were merged.</p>",
      "date_published": "2026-07-29T09:00:00.000Z",
      "tags": [
        "normalisation",
        "titles",
        "engineering"
      ]
    },
    {
      "id": "https://hiringindex.org/blog/reading-hiring-momentum",
      "url": "https://hiringindex.org/blog/reading-hiring-momentum",
      "title": "Reading hiring momentum without fooling yourself",
      "summary": "A 30% jump in postings usually means something changed in the pipeline, not in the economy.",
      "content_html": "<p>Change over time is the most-read and least-reliable number in job posting data. Here is how to tell a real movement from an artefact.</p>\n<h2>Four false alarms</h2>\n<ol><li><strong>A new source.</strong> Adding a large employer or a board to an index looks exactly like a hiring surge. Any provider that changes coverage without saying so will produce spurious momentum.</li><li><strong>Seasonality.</strong> January and September are structurally busy; late December is structurally dead. Comparing to the previous month rather than the same month last year turns the calendar into a finding.</li><li><strong>A single large employer.</strong> One retailer opening seasonal roles can move a national count for a category by double digits.</li><li><strong>Republishing.</strong> A mass refresh of old listings creates a wave of new postings without a single new role.</li></ol>\n<h2>Three checks that take a minute</h2>\n<ul><li><strong>Look at the age distribution.</strong> A genuine surge is made of young listings. A refresh artefact is made of listings whose content is not new.</li><li><strong>Look at the employer concentration.</strong> If the top employer accounts for most of the change, it is one company's decision, not a market.</li><li><strong>Compare year on year.</strong> It removes seasonality at the cost of sensitivity, which is the right trade for most questions.</li></ul>\n<h2>What movement is worth acting on</h2>\n<p>Sustained, broad and young: a change that holds for more than one period, is spread across employers, and shows up in fresh listings rather than recycled ones. Everything else is worth watching and not worth a decision.</p>\n<p><em>We show the thirty-day change alongside the age distribution and the employer concentration on every cohort page, so these three checks are on the same screen rather than requiring three queries.</em></p>",
      "date_published": "2026-07-27T09:00:00.000Z",
      "tags": [
        "analysis",
        "momentum",
        "labour market"
      ]
    },
    {
      "id": "https://hiringindex.org/blog/how-long-to-wait-after-applying",
      "url": "https://hiringindex.org/blog/how-long-to-wait-after-applying",
      "title": "How long to wait after applying, according to the listings themselves",
      "summary": "You cannot see inside a company's process. You can see how long its postings usually stay open.",
      "content_html": "<p>The standard advice — follow up after a week, assume no news is bad news after two — is guesswork dressed as a rule. There is a better proxy, and it is public.</p>\n<h2>Use the posting, not the process</h2>\n<p>You cannot observe a hiring pipeline from outside. You can observe how long that employer's listings usually stay up before disappearing. If a company's postings typically vanish after eighteen days, a role you applied to on day two is probably being decided inside three weeks. If its postings sit for four months, your application is entering a queue, not a process.</p>\n<h2>What typical looks like</h2>\n<p>Median posting age varies enormously by role and market, which is precisely why an average is useless. Some patterns hold:</p>\n<ul><li>High-volume roles — retail, hospitality, support, driving — turn over fastest. The listing is often gone within a fortnight because it was filled or the requisition expired.</li><li>Specialist and senior roles sit longest. Two to three months is normal and says nothing bad about the employer.</li><li>Roles at large firms sit longer than roles at small ones, because internal process is longer, not because interest is lower.</li></ul>\n<h2>A practical reading</h2>\n\n<h2>The uncomfortable part</h2>\n<p>None of this tells you whether anyone read your application. It tells you whether the requisition is alive. Those are different questions, and only the second one is visible from outside — which is exactly why we measure it and do not claim the first.</p>",
      "date_published": "2026-07-24T09:00:00.000Z",
      "tags": [
        "job search",
        "posting age",
        "process"
      ]
    },
    {
      "id": "https://hiringindex.org/blog/job-search-ux-what-helps",
      "url": "https://hiringindex.org/blog/job-search-ux-what-helps",
      "title": "Job search interfaces: what helps and what is decoration",
      "summary": "People filter by outcome, not by taxonomy. Most job search UI is organised the wrong way round.",
      "content_html": "<p>Job search interfaces converge on the same layout: a keyword box, a location box, and a column of filters copied from whatever the database happens to store. That column is where most of the wasted effort goes.</p>\n<h2>What people actually filter on</h2>\n<ol><li><strong>Can I do this from where I live?</strong> Remote, hybrid, on-site — the first cut almost everyone makes.</li><li><strong>Does it pay enough?</strong> A floor, not a range.</li><li><strong>Am I plausible for it?</strong> Seniority, expressed in words people use about themselves.</li><li><strong>Is it still open?</strong> Rarely offered as a filter, and near the top of what people want.</li></ol>\n<p>Notice what is missing: industry, company size, contract type. They exist in every filter column and get used by a small minority.</p>\n<h2>Filters that are worth their space</h2>\n\n<h2>Three things that help more than another filter</h2>\n<ul><li><strong>Show the age of every listing.</strong> One number per result, and it answers the question people are already asking.</li><li><strong>Put the state in the URL.</strong> A search someone can send to a friend is worth more than one they have to describe.</li><li><strong>Make the empty state useful.</strong> \"No results\" is a dead end. \"No results — try removing the salary floor, which excluded 340 roles\" is a next step.</li></ul>\n<h2>The result card</h2>\n<p>Four things earn their place: title, employer, location and pay. Everything else competes with them. If pay is unknown, say so plainly rather than hiding the row — an absent field reads as an oversight, an explicit \"not disclosed\" reads as honesty and is also information.</p>",
      "date_published": "2026-07-22T09:00:00.000Z",
      "tags": [
        "ux",
        "product",
        "job boards"
      ]
    },
    {
      "id": "https://hiringindex.org/blog/remote-work-what-the-postings-say",
      "url": "https://hiringindex.org/blog/remote-work-what-the-postings-say",
      "title": "Remote work: what the postings actually say",
      "summary": "Advertised remote share is a measurement of what employers are willing to write down, which is not the same as how people work.",
      "content_html": "<p>Remote share is the most quoted statistic in job posting data and the most misread. It measures a specific, narrow thing: the proportion of advertised roles that describe themselves as remote.</p>\n<h2>Three things it is not</h2>\n<ul><li><strong>Not the share of people working remotely.</strong> Existing staff are invisible to posting data. A company can be fully remote internally and advertise everything as hybrid.</li><li><strong>Not a measure of flexibility.</strong> \"Hybrid\" covers one day a week and four. The word is doing a lot of work.</li><li><strong>Not comparable across sources.</strong> Some providers infer remote from the description; others use a structured field. Inference finds more and is wrong more often.</li></ul>\n<h2>What it is good for</h2>\n<p>Direction, within one source, over time. If the advertised remote share for a role falls from 40% to 25% over a year in the same index, employers have changed what they are willing to commit to in writing. That is a real finding even though the underlying working patterns may have moved less.</p>\n<h2>The pattern that holds everywhere</h2>\n<p>Remote share varies far more by role than by country. Roles whose output is a file — engineering, design, writing, analysis — are advertised remote at multiples of the rate for roles that touch people, inventory or equipment. National differences are mostly a composition effect: countries with more of the first kind show higher headline remote shares.</p>\n<p>Which means a national remote-share comparison is usually an industry-mix comparison wearing a disguise. Compare within a role.</p>\n<h2>The salary question</h2>\n<p>Remote roles sometimes advertise lower ranges than on-site equivalents in the same city, and sometimes higher. Both happen, and the direction depends on whether the employer is anchoring to their own market or competing in a national one. Anyone quoting a single \"remote pay penalty\" figure has averaged over two opposite behaviours.</p>",
      "date_published": "2026-07-19T09:00:00.000Z",
      "tags": [
        "remote work",
        "labour market",
        "analysis"
      ]
    },
    {
      "id": "https://hiringindex.org/blog/what-a-median-salary-hides",
      "url": "https://hiringindex.org/blog/what-a-median-salary-hides",
      "title": "What a median salary tells you, and what it hides",
      "summary": "A single number describes a distribution badly. Four numbers describe it well enough to act on.",
      "content_html": "<p>A median is the middle value: half the postings are below, half above. It is the right summary for salary data because it survives outliers — one absurd executive package does not drag it around the way an average would.</p>\n<p>It is still one number describing a shape, and the shape matters.</p>\n<h2>Same median, different markets</h2>\n<p>Consider two cohorts with an identical median of $130,000:</p>\n\n<p>In A the role is standardised: seniority barely moves pay, and negotiation has little room. In B the title spans junior to staff, and your outcome depends almost entirely on which end you land in. Same median, completely different advice.</p>\n<h2>The four numbers worth reading</h2>\n<ul><li><strong>25th percentile</strong> — what the bottom of the market looks like. Below it, you are being underpaid relative to the advertised market.</li><li><strong>Median</strong> — the centre. A reasonable target if your experience is typical for the cohort.</li><li><strong>75th percentile</strong> — reachable with strong, relevant experience.</li><li><strong>90th</strong> — usually a different job wearing the same title.</li></ul>\n<h2>Three ways a median misleads</h2>\n<ol><li><strong>Selection.</strong> It covers only postings that published a range. If disclosure is 15%, the median describes those 15%.</li><li><strong>Level mixing.</strong> \"Engineer\" spans a factor of three. Split by seniority before drawing conclusions.</li><li><strong>Geographic mixing.</strong> A national median blends metros with very different costs. Narrow to a city.</li></ol>\n<h2>How we present it</h2>\n<p>Every cohort page shows the histogram, the four percentiles, the split by seniority, and the disclosure rate. That is more than one number because one number would let you reach a confident wrong conclusion, and the whole point of the product is not doing that.</p>",
      "date_published": "2026-07-17T09:00:00.000Z",
      "tags": [
        "salary",
        "statistics",
        "method"
      ]
    },
    {
      "id": "https://hiringindex.org/blog/what-we-do-not-know",
      "url": "https://hiringindex.org/blog/what-we-do-not-know",
      "title": "What we do not know",
      "summary": "A list of the gaps in our own data, published because you will find them anyway.",
      "content_html": "<p>Every data product has holes. Most vendors leave you to discover theirs. Here are ours, so you can decide whether they matter for what you are doing.</p>\n<h2>Coverage</h2>\n<ul><li>We do not cover every employer. Sources that restrict automated access are absent, and we do not estimate around the gap.</li><li>Coverage is uneven by country and by employer size. Smaller employers who post only to a local board are underrepresented.</li><li>A market where we have thin coverage will produce a plausible-looking median computed from too few postings. The counts are always shown so you can judge.</li></ul>\n<h2>Salary</h2>\n<ul><li>Only postings that publish a range are included in any salary figure.</li><li>Disclosure is not random: it skews to regulated markets and larger employers. Medians inherit that skew.</li><li>We do not convert currencies. A cross-border comparison of our medians compares tax systems as much as pay.</li></ul>\n<h2>Age and freshness</h2>\n<ul><li>Age is computed from the publication date, not from continuous observation.</li><li>Republishing resets it, and we cannot always detect that.</li><li>A listing being visible does not prove the role is open.</li></ul>\n<h2>Titles and categories</h2>\n<ul><li>Grouping by title is a judgement call. Two people would draw the lines differently.</li><li>Composite roles land in one bucket.</li><li>New role names take time to appear in any grouping, including ours.</li></ul>\n<h2>Intent</h2>\n<p>We cannot see whether an employer means to hire. Nothing we publish should be read as a claim about intent, and where our pages discuss stale listings they say \"signal\", not \"verdict\", deliberately.</p>\n<h2>Why publish this</h2>\n<p>Two reasons. A buyer who finds a limitation on their own concludes it was hidden; a buyer who reads it here concludes it was measured. And writing the list down keeps us honest about which gaps we have accepted and which we are still working on.</p>",
      "date_published": "2026-07-15T09:00:00.000Z",
      "tags": [
        "method",
        "transparency",
        "limitations"
      ]
    }
  ]
}