Non-Obvious Candidate Signals That Predict Engineering Performance

Coinbase ran an internal analysis comparing two different rounds of its engineering interview loop and found an 84% correlation between them.
That finding is the clearest evidence of a problem that shows up across most engineering hiring loops: they produce false positives and false negatives at the same time. The false positive is familiar to anyone who has sat through a whiteboard round. Candidates who put in enough study time can recite the canonical answer for a URL shortener, a distributed queue, or a caching layer, and the interview rewards the recitation. Many of these are the same engineers best suited to a workplace increasingly built around directing and correcting AI rather than typing out algorithms from memory.
The same pattern appears in both failures: candidates have to get good at interview preparation to land the job, but it has no bearing on whether they can do the job once they're in it. An 84% correlation between two rounds shows a redundant process, run twice on the same narrow slice of a candidate's ability, while the engineering judgment that actually determines performance on the job goes unmeasured.
How AI-generated code changed engineering work and hiring requirements
That mismatch would be a problem on its own. What makes it urgent is how fast the actual content of engineering work has moved out from under it. By Q4 2025, AI-generated code had crossed the halfway mark of everything merged, while human-authored code fell sharply over the same stretch. Human review, meanwhile, stayed close to universal. Engineers didn't stop reviewing work. They stopped being the ones who typed most of it.
That changes what the job is made of day to day. A Coinbase engineering leader put the shift this way: when the cost of building goes to zero, the cost of identifying what to build, verifying it's correct, and getting it out safely becomes the limiting factor.
That sentence should sit uncomfortably next to most engineering interview loops still in use, because most of them were built for a world where writing the code was the hard part. A loop built to test whether someone can write a clean for-loop from memory is testing a skill that AI now performs by default. The candidate who passes that loop by demonstrating strong recall may be the one least prepared for the job as it exists now, because the job has moved from authorship to judgment, and recall tests don't measure judgment. The gap between what the old loop checks and what the 2026 job demands didn't close. It widened, and it widened inside about a year.
How AI fluency works as a hiring signal
If the job now runs on directing and correcting AI, then "uses AI well" becomes a hiring signal worth paying close attention to. But it isn't one trait a candidate either has or doesn't. It splits into at least three distinct levels, and treating them as interchangeable just recreates the same false-positive problem the old loop had, this time dressed up as AI fluency.
The first tier is the fluent reviewer. That combination, speed from the model plus a sharp, consistent check on its output, fits senior IC roles best, where an architectural mistake is costly and judgment affects outcomes more than raw output does.
The second tier is the productive user. It's a better fit for a different job: mid-level product engineering on a team where shipping speed is the actual constraint and the cost of a smaller mistake is tolerable.
The third tier is the AI-native builder. This engineer designs systems around AI agents from the start and treats prompt engineering as a real part of the engineering surface. That tier fits AI and ML platform roles specifically, where the product itself is built on orchestrating models.
None of these three map cleanly to seniority. The question a hiring loop needs to ask is which tier the role actually needs, and then whether the candidate in front of them demonstrates it.
The sharpest version of that signal, across all three tiers, is pushback. A candidate who pastes whatever the model produces is showing the opposite, no matter how fast the output came.
Coinbase rebuilt its frontend coding interview around this idea directly. The new versions give candidates full AI access on realistic problems built so the signal comes from how a candidate uses the tool: the quality of their prompts, how carefully they check the output, how fast they spot an error, how they adjust after a first pass that didn't land. Producing a good solution, and knowing when to stop trusting the model and start checking it, was not. Candidates who passed this AI-assisted assessment went on to clear Coinbase's onsite interviews at a meaningfully higher rate than candidates who passed the old version, which is a strong indication that the new version is measuring something closer to the actual job.
Proof-of-work signals that surface engineering judgment without relying on credentials or interview performance
AI fluency is one strong signal, but it's not the only place engineering judgment becomes visible. The work a candidate has already done tells you more than work they perform under the artificial pressure of an interview room. Github history is one place to look, though not in the way most people glance at it. Documentation and tests a candidate chose to write, without being told to, say something about how they think about reliability and about communicating with teammates, not just about whether they can produce working code. Even attendance at developer meetups or steady participation in technical communities says something: it points to time spent on the craft that has nothing to do with chasing a job offer.
One question does more diagnostic work than almost anything else available: ask a candidate what percentage of a given pull request was AI-generated, and what they personally changed. The answer separates engineers who are still growing their own judgment as the tools shift under them from engineers who've quietly stopped growing the moment the model started doing the typing. It's a direct, concrete way to find out whether someone is directing the work or just forwarding it.
System design interviews still matter here, but only the ones built around architecture, tradeoffs, and failure modes rather than a memorized pattern. That kind of question reveals whether a candidate can design something a model will later help build, which is a different skill than reciting the standard answer to a familiar prompt. The interview becomes a compressed version of the job instead of a separate skill a candidate has to learn on top of it.
None of this requires a wide résumé. An engineer with one clearly demonstrated strength, visible in real contributions, is often a stronger bet than one with ten shallow ones spread across a dozen frameworks, even though the second résumé reads as more impressive at a glance. The obvious objection is that leaning on public work samples favors people with the free time to build a visible portfolio, and that's a fair concern. An engineer who left careful, specific feedback on someone else's pull request is demonstrating judgment, whether or not they had the spare hours to ship a personal app on the side.
Why non-obvious candidates get eliminated before these signals can be seen
None of these signals matter if the candidate never reaches a stage where someone looks for them. For AI engineering roles specifically, the large majority of senior candidates work at frontier labs, hyperscalers, AI-focused scale-ups, or specialist boutiques, and they are not actively browsing job boards. A standard inbound posting reaches only a small slice of that pool. The problem sits upstream of the interview entirely: sourcing built around matching job titles and keywords pulls in roughly the same kind of candidate the job description already describes, which is exactly the shallow proxy the rest of this piece has been arguing against.
Part of the trouble is definitional. AI engineer, ML engineer, data scientist, AI research scientist, and AI software engineer get used almost interchangeably in job postings and sourcing tools, but the candidate pools behind those titles only partly overlap. Treating them as the same thing from the start means the pool a recruiter builds is structurally wrong before a single résumé gets a human look.
A gap in most loops appears around reliability. The engineer worth hiring for most of these roles is the one who treats keeping a non-deterministic system reliable as the actual job, not an afterthought. A loop that never asks about monitoring, failure handling, or what happens when a model's output drifts will never find that person, no matter how well-designed the rest of the interview is.
Putting these pieces together leaves a hiring team with a pool that looks thinner than the market actually is. The honest response is to fix sourcing and screening, but the common response is to assume the talent simply isn't out there, lean harder on prestigious employers and familiar credentials as a shortcut, and repeat the mistake that built the shallow pool.
What calibration requires for better signals
Better signals are only useful to a team that already knows what it's looking for. The starting point has to be defining the two or three non-negotiable competencies for a specific role before a single interview question gets written. A backend engineer who needs to design scalable APIs well, a frontend developer whose main job is UI performance, and a DevOps specialist focused on cloud infrastructure automation are three different definitions of success, and each one calls for a different set of signals to check for.
Those competencies only become useful once they're translated into behaviors an interviewer can actually watch for in real time, rather than left as abstract traits on a job description. This is where a lot of otherwise promising tools fall apart in practice. If a sourcing or screening tool hands a hiring manager a pool of candidates with some kind of score attached, and the manager doesn't understand what that score is actually measuring, the natural response is to throw out the whole pipeline and go back to doing it by hand. That looks like a failure of the technology. It's usually a failure of calibration: nobody translated what "exceptional" means for this specific role into something the tool, or the interviewer, could check for.
The same logic applies to AI-inclusive live exercises, where a candidate talks through their thinking out loud: why they accepted or rejected a given AI suggestion, how they checked the output before trusting it, how they handled an edge case the model missed. That exercise only produces a real signal if the interviewer walked in already knowing what separates a strong answer from a weak one. Without that groundwork, the exercise just produces more conversation, not more clarity.
Calibration isn't something to settle once and then leave alone. Every interview and every hire either sharpens the definition of what "exceptional" means for a role or quietly lets it drift, and any change to that standard should be visible to the team and approved by a human, not something that shifts on its own between hiring cycles. It also has to stay honest about the market. The 2026 hiring market rewards specific, legible specialties: LLM application work, MLOps, infrastructure at scale, named technical domains. Calibration has to include a realistic read of what that market will actually produce for a role as defined, or the standard being calibrated toward may not exist in the pool being sourced from.
A unified hiring system for capturing non-obvious signals, keeping humans in the decision seat, and avoiding handoff failures
Even a well-calibrated set of signals gets lost if the process carrying them is broken into disconnected pieces, each handled by a different tool or a different person with only partial visibility into what came before. That fragmentation is itself a reason non-obvious candidates disappear, separate from any problem with the signals themselves.
A calibration conversation between a recruiter and a hiring manager might produce a precise, well-reasoned definition of what exceptional looks like for a role. A strong signal from a technical screen, maybe a sharp answer about handling a pull request, or a thoughtful walkthrough of a system design tradeoff, often gets reduced to a short summary in a debrief note that the next interviewer reads days later, with no access to the original conversation or artifact. And speed compounds the damage: top AI engineering candidates routinely take competing offers within three weeks, while the industry-average time-to-hire runs to 47 days. A candidate who cleared every real bar can still be lost simply because the process was between handoffs when a competing offer landed.
Fixing this means giving AI the parts of the process it's genuinely suited to while people stay in the decision: outbound sourcing aimed at passive candidates using non-credential signals instead of keyword matching, first-pass screening measured against the specific behavioral indicators a team has already calibrated, and summarization of technical assessments so that context, the actual prompts used, the specific error caught, the tradeoff weighed, travels with the candidate instead of getting flattened into a one-line note. Humans stay in the seat where judgment about a person actually belongs: deciding what exceptional means for the role, reading the original signal rather than a summary of it, and making the call. The system's job is to make sure that signal survives the trip from first contact to final decision, intact enough for the humans doing the choosing to still recognize the candidate they calibrated for.


