ROI Case File No.566: 'Accuracy Went Up, but the Same People Kept Passing'
![]()
Accuracy Went Up, but the Same People Kept Passing
Chapter 1: Accuracy Rose, but Hiring Grew Uniform
"I want to reshape our new-graduate hiring AI. With the current tool, I feel a ceiling."
Kohei Senda, HR director at TechSolutions, recounted the background. "For several years we've used Mynavi's AI tool to read entry sheets and score pass likelihood from past selection data. Accuracy did go up. But every year, the same kind of entry sheets pass. The talent we hire has grown uniform."
"Accuracy rose, but a problem emerged," Claude asked.
"That's it," Senda answered. "The AI accurately picks people who resemble past hires. So accuracy is high. But 'resembles the past' is all it means, and diversity was lost. The lineup became the same faces."
"Is there a change on the students' side, too?" I asked, to confirm.
"There is," Senda answered. "More students write their entry sheets with generative AI, and even human eyes struggle to tell. The current tool can't handle that. So I want to drop the existing tool and consider a new form of AI use."
"If raising accuracy makes it more uniform, you need to rebuild the evaluation criteria at the root," I replied. "Let's break this down with DESC."
Chapter 2: DESC Asks—Rebuild Through Describe, Express, Suggest, Consequence
"This case calls for DESC."
Claude wrote on the whiteboard: "Describe, Express, Suggest, Consequence."
"DESC—Describe, Express, Suggest, Consequence; fact, impact, proposal, result—frames a problem starting from objective fact, shows the impact, makes a proposal, and foresees the result, in that order," I explained. "The key is to start from fact, not impression. Drop the impression of 'the same people' onto the fact of the data. What is happening, what it produces, how to change it, what results—assemble them in order, and you can rebuild the evaluation criteria from the root."
"First, let's measure the current cost," Gemini said, opening ROI Polygraph. The data Senda provided was entered.
"Here is the monthly cost," Gemini read out. "Labor for selection work, including entry-sheet checks and generative-AI detection: 150 hours per month on average, at ¥3,800 per hour, ¥570,000 per month. Opportunity loss from lost diversity and organizational rigidity due to uniform hiring: ¥420,000 per month. Risk of degraded evaluation accuracy from the difficulty of detecting generative-AI-written entry sheets: ¥360,000 per month. Missed promising talent from over-reliance on past data: ¥380,000 per month. Early-turnover risk from the gap between the aptitude test and the real person: ¥320,000 per month. A total of ¥2,050,000 per month. Roughly ¥24.60 million per year."
Senda stared at the figures. "I was only looking at the labor of selection. Add the cost of uniformity and missed talent themselves, and it comes to this much."
"Then let's design it with DESC," I continued.
[Describe—Capture the bias of the past data objectively]
"First, capture the fact," Claude said. "Analyze the past selection data in detail and grasp objectively what kind of entry sheets have passed. Drop the impression of 'the same kind of person' onto the fact of the data. Start from fact, and it doesn't become argument-by-feeling."
[Express—Show what uniformity produces]
"Next, show the impact," Gemini continued. "Pass only people who resemble the past, and diversity is lost and the organization rigidifies. Fail to see through generative-AI-written entry sheets, and the evaluation itself wavers. Show clearly what the bias produces."
[Suggest—Rebuild the criteria and add detection]
"After impact, propose," I continued. "Redefine the evaluation criteria and stop the past-data bias. Develop a new algorithm to detect generative-AI-written entry sheets. Further, build in a personality-aptitude test to minimize the gap between the ideal and the real self. Rebuild the criteria at the root."
[Consequence—Foresee hiring where diverse talent passes]
"Finally, foresee the result," Claude continued. "Rebuild the criteria and promising talent that doesn't resemble the past also passes. Generative-AI-written entry sheets can be detected, and the evaluation doesn't waver. The aptitude test captures the real person, and post-hire mismatch drops. Draw the picture beyond the rebuild first."
[Estimating the payback]
"Let's run the numbers on ROI Proposal Generator," Gemini proposed.
- Initial cost: redefining evaluation criteria, developing the generative-AI detection algorithm, personality-aptitude-test integration, building the selection-support system, and operational design—¥5.2 million total
- Monthly cost: system operation and model updates combined, ¥210,000
- Monthly savings: streamlining selection work = ¥400,000 (assuming a 70% reduction); diversifying talent by dissolving uniformity = ¥400,000; securing accuracy via generative-AI detection = ¥340,000; reducing missed talent and curbing mismatch = ¥340,000; totaling ¥1,480,000 per month
- Net monthly savings: ¥1,480,000 − ¥210,000 = ¥1,270,000 per month
- Payback period: ¥5.2 million ÷ ¥1.27 million = about 4.1 months
"Just over four months to recoup," Gemini summarized. "What works is not merely raising accuracy, but rebuilding the evaluation criteria from the root. Raise accuracy by resembling past data and uniformity advances. Recapture from fact and recombine the criteria, and diverse talent also passes. The investment doesn't miss."
Senda checked the figures. "I thought raising accuracy would make it better. Unless you rebuild the criteria themselves, it's the same people."
"DESC is a tool for rebuilding evaluation criteria from fact," I replied.
Chapter 3: An Implementation Plan That Rebuilds the Criteria
"Let me lay out the approach," I said, standing at the whiteboard.
"Month one—analyze past selection data and grasp the fact of pass tendencies. Month two—assess the impact of uniformity and redefine the evaluation criteria. Months three and four—develop the generative-AI detection algorithm and build the selection-support system. Month five—integrate the personality-aptitude test and design operations. Month six—trial operation and effect verification (measuring diversity and accuracy). Month seven onward—continuously review the evaluation criteria and embed them into next year's hiring."
"Can we keep accuracy and still get diversity?" Senda asked.
"You can," Claude replied. "They can't coexist because you get accuracy by resembling past data. Recapture the fact with DESC and recombine the evaluation criteria themselves. Evaluate not by closeness to past hires but by factors that lead to success. Change the criteria, and diverse talent passes while accuracy holds. The 'same people only' ends."
Senda took notes. "Before chasing accuracy, rebuild the criteria from fact. Now I see the order."
Chapter 4: The Day the Same Faces Changed
Ten months later, a report arrived from Senda.
After the criteria were rebuilt, the talent hired changed. "From the same kind of person only, students of diverse backgrounds began to pass. A new wind entered the organization," Senda wrote.
Generative-AI-written entry sheets could be seen through, too. The detection algorithm handled judgments that were hard for the human eye. "AI-written entry sheets could be detected, and the evaluation stopped wavering," the report read.
The biggest change showed in how evaluation was grasped. From measuring by resemblance to the past, to measuring by factors that lead to success. "We were picking people close to past hires. Once we rebuilt the criteria from fact, we could pick up the promising even when they didn't resemble," Senda wrote.
Post-hire mismatch dropped as well. The aptitude test curbed the gap between the ideal and the real. "The 'didn't fit after joining' decreased," the report read.
As a side effect, the way hiring was viewed changed. Not merely chasing accuracy, but questioning the criteria, took root. "We stopped 'just raise the AI's accuracy.' We started asking what we should evaluate," Senda wrote.
At the end of Senda's report was this: "I thought the new-grad hiring struggle was evaluation accuracy. But the real problem was that raising accuracy made it the same people who resembled the past, and hiring grew uniform. The moment we rebuilt the evaluation criteria from fact with DESC, the road to diverse talent came into view. Before chasing accuracy, questioning the criteria came first."
The day a company where accuracy rose but the same people kept passing became a company that could choose diverse talent, hiring AI had shifted from the pursuit of accuracy to a design that rebuilds the criteria through describe, express, suggest, and consequence.
"Hiring-AI requests usually arrive as 'I want to raise evaluation accuracy.' But before chasing accuracy, there's a question to ask: what is that accuracy resembling? What DESC asks is describe, express, suggest, and consequence. Raise accuracy by resembling past data and hiring grows uniform. Rebuild the criteria from fact, and diverse talent also passes. The day a company where accuracy rose but it was the same people could change its lineup, what changed was not the AI's accuracy but the very perspective that rebuilds evaluation criteria from fact."
Related Files
Tools Used
- ROI Polygraph — Visualizing selection-work labor, opportunity loss from uniformity, and missed promising talent
- ROI Proposal Generator — Payback simulation for new-graduate hiring AI, starting from rebuilding the evaluation criteria from fact