Note Wisdom
Most hiring tools—resumes, unstructured interviews, high-stakes personality tests—produce validity coefficients near zero when tested against actual job performance. Applying James Randi’s skeptical standard of evidence reveals that organizations continue using ineffective tools due to psychological bias and structural inertia, not empirical support. Evidence-based alternatives exist and should replace intuition-driven practices.
Walk into almost any corporate HR department and you will witness a ritual that would make James Randi smile—not because it works, but because everyone pretends it does. Candidates submit resumes that are essentially creative writing exercises. Recruiters scan them for keywords that have zero correlation with future performance. Interviewers ask “Tell me about yourself” and nod thoughtfully, as if the answer reveals something predictive. Then everyone wonders why turnover remains high and performance reviews disappoint.
Randi, the legendary skeptic who publicly swallowed an entire bottle of homeopathic sleeping pills on a TED stage to prove they contained nothing active, built his career on one simple principle: test the claim. If a product claims to cure insomnia, ingest it and see what happens. If a psychic claims to read minds, design a double-blind experiment and watch them fail. The homeopathy industry thrives because people want to believe—not because the evidence supports it.
Personnel selection has a homeopathy problem of its own. Organizations pour billions into recruitment processes built on tools that, when actually tested against the one criterion that matters—subsequent job performance—produce validity coefficients indistinguishable from random noise. The recruitment industry, like the alternative medicine industry, survives on the same mechanism Randi diagnosed: audience assumptions. We assume resumes tell us something. We assume unstructured interviews reveal character. We assume a candidate who speaks well will perform well. We assume these things because they feel right, not because the data says so.
This article applies Randi’s skeptical research paradigm—parameter disassembly, defect matching logic, and root cause tracing—to the personnel selection tools organizations use every day. The goal is not to declare all hiring practices worthless. The goal is to separate what actually predicts performance from what merely feels like it should.
Before we can identify which tools work, we need a unit of measurement. In personnel selection psychology, that unit is the validity coefficient—a correlation coefficient ranging from zero to one that expresses how well a selection tool predicts subsequent job performance. A coefficient of .30 means the tool accounts for roughly 9 percent of the variance in performance. A coefficient of .50 accounts for 25 percent. At scale, across hundreds of hires, that difference translates into millions of dollars in productivity.
The field’s foundational meta-analyses have given us clear benchmarks. General cognitive ability tests consistently produce operational validities in the .40 to .60 range. Structured interviews, when properly designed with standardized questions and scoring rubrics, achieve coefficients around .51. Work sample tests—where candidates actually perform tasks resembling the job—fall in the .30 to .40 range.
These are the tools that pass Randi’s test. They have been subjected to meta-analytic scrutiny across thousands of studies and millions of employees. Their effects replicate across industries, countries, and job types.
But these are not the tools most organizations actually use.
Let us start with the most ubiquitous selection tool in existence. Every job application requires one. Recruiters spend hours reading them. And yet, a 2025 Academy of Management study found that resume assessments produce validity coefficients between .04 and .07—essentially zero. Resume evaluations were “not significantly related to job performance”. Another study put the correlation at .05.
Think about what this means. You are better off flipping a coin than reading a resume if your goal is to predict who will perform well. Resumes are exercises in self-presentation, not samples of work behavior. They tell you what candidates want you to know, filtered through whatever narrative they believe will get them an interview. They are the homeopathic remedy of hiring—diluted to the point of irrelevance, yet consumed in massive doses because everyone assumes they must contain something active.
The unstructured interview—the casual conversation where interviewers ask whatever comes to mind—produces validity coefficients around .20. Even the more generous meta-analytic estimates put it at .38. Either way, it barely outperforms chance.
Why? Because unstructured interviews are riddled with noise. Different interviewers ask different questions. Different candidates receive different follow-ups. The same candidate can receive wildly different ratings from two interviewers who witnessed the exact same conversation. Interviewers form impressions in the first few minutes and spend the rest of the conversation confirming their initial bias. They ask questions that feel relevant but have no demonstrated link to job performance.
Randi would recognize this immediately. The unstructured interview depends entirely on the interviewer’s assumption that they can accurately judge character and capability from a thirty-minute conversation. But as Randi demonstrated with his empty glasses and beard trimmer, human beings are terrible at detecting when they are being deceived—especially when they want to believe.
Personality assessments present a more subtle case. In low-stakes settings—employee development, academic research—they show modest validity. Conscientiousness, for example, correlates with performance at around .22 to .27. But in high-stakes hiring contexts, where applicants have every incentive to present themselves in the most favorable light, validity plummets.
A 2025 meta-analysis compared personality test validity across low-stakes and high-stakes settings. The findings were stark: in matched samples, low-stakes validity was .27, while high-stakes validity was .12—a 125 percent difference. The researchers concluded that “faking substantially reduces personality test validity in selection contexts”.
This is the homeopathy problem in action. The tool can work under ideal conditions. But the conditions under which it is actually used—job applicants highly motivated to fake—render it largely ineffective. Organizations that rely on off-the-shelf personality tests for hiring are administering a remedy that only works in the laboratory, not in the real-world context where it matters.
If these tools do not work, why does every organization still use them?
The answer lies in the same psychological mechanisms Randi spent his career exposing. Human beings are pattern-seeking creatures who prefer comfortable stories over uncomfortable data. A resume feels informative. A conversation feels revealing. A personality test feels scientific. These feelings persist regardless of what the validity coefficients say.
There is also a structural incentive problem. HR professionals are evaluated on process efficiency—how quickly they fill roles, how many candidates they screen—not on the predictive accuracy of their tools. No one gets fired for using a resume. No one gets sued for conducting unstructured interviews. The tools persist because the cost of changing them is visible, while the cost of keeping them is invisible.
Randi would call this what it is: fraud by omission. Not deliberate fraud, perhaps, but fraud nonetheless—the continued use of tools that have been empirically demonstrated not to work, justified by nothing more than tradition and convenience.
The solution is not to abandon all selection tools. The solution is to use only the tools that have survived meta-analytic scrutiny, and to combine them in ways that maximize predictive power while minimizing bias.
Stop assuming your current tools work. Test them. If you are using resumes, unstructured interviews, or personality tests in high-stakes settings, you are almost certainly wasting money and missing talent. Run a concurrent or predictive validity study on your own organization’s data. Correlate selection scores with six-month or twelve-month performance ratings. If the coefficient is below .20, the tool is not doing its job.
General cognitive ability tests remain the single most valid predictor of job performance across virtually all occupations. They are not without controversy—they produce adverse impact against certain demographic groups—but excluding them “generally has little to no effect on validity, while substantially decreasing adverse impact,” according to a 2024 Journal of Applied Psychology meta-analysis. The trade-off is real, but the solution is not to abandon ability testing; it is to combine it with other valid, lower-impact tools.
Structured interviews, when designed around job-relevant competencies with standardized questions and scoring rubrics, achieve validity coefficients around .51. They also tend to reduce adverse impact compared to unstructured formats. Work sample tests, where candidates perform actual job tasks, produce validities in the .30 to .40 range and are highly resistant to faking.
Situational judgment tests—where candidates respond to hypothetical work scenarios—show useful validity around .34 and add incremental prediction beyond cognitive ability and personality. Assessment centers, despite their cost, produce validities from .28 to .63 depending on design. Integrity tests, while controversial, have demonstrated criterion-related validity in multiple meta-analyses.
This is the step most organizations skip. They purchase a test, implement it, and never check whether it actually predicts performance in their context. But validity is not a property of the test alone—it is a property of the test in a specific context with a specific population.
The 2025 meta-analysis on personality test validity made this point explicitly: “Practitioners of personnel selection should validate or revalidate assessments under high-stakes conditions”. Low-stakes validity evidence is provisional at best. If you are using a tool for hiring, validate it with hiring data.
James Randi’s contribution to science was not just debunking individual frauds. It was establishing a standard of evidence—a refusal to accept claims without data, a commitment to designing tests that could actually falsify the claim being made.
Personnel selection needs the same standard. The question is not “Does this tool feel useful?” or “Do candidates like it?” or “Is it popular in my industry?” The only question that matters is: Does it predict job performance?
If the answer is no—and for resumes, unstructured interviews, and high-stakes personality tests, the meta-analytic answer is unequivocally no—then continuing to use the tool is not just inefficient. It is, in Randi’s terms, a form of quackery. It exploits the credulity of candidates and the assumptions of decision-makers, delivering nothing of value while consuming time and resources that could be spent on tools that actually work.
The good news is that we know what works. Cognitive ability tests, structured interviews, work samples, and properly validated situational judgments all clear the bar. The challenge is not scientific—it is organizational. It is convincing leaders to abandon the comfortable fictions they have relied on for decades and adopt evidence-based practices that, while sometimes less intuitively satisfying, actually predict who will succeed.
That is the real lesson from Randi’s sleeping pills. The homeopathic remedy felt legitimate. It came in a proper bottle with proper instructions. It looked like medicine. But when tested—really tested—it contained nothing. Most of your hiring process is the same. It looks like selection. It feels like selection. But when tested against the only criterion that matters, it predicts nothing.
Time to swallow the whole bottle and see what happens.
Source Reference Link: https://www.ted.com/talks/james_randi_homeopathy_quackery_and_fraud
Link Brief: Legendary skeptic James Randi delivers a fierce critique of homeopathy and pseudoscientific fraud. He publicly takes an entire bottle of homeopathic sleeping pills with no adverse effects, proving these diluted products contain no active ingredients, and exposes how quacks exploit public psychological credulity for profit. This article applies Randi’s skeptical evidentiary standard to personnel selection, arguing that most common hiring tools fail the same empirical test.
Content Disclaimer: This article is for general reference only and does not constitute professional R&D guidance, production process advice or quality certification. All material performance data has specific test premises; readers should verify parameters against actual equipment and working conditions.

