Do you still evaluate engineers based on the code they produce? If so, you're lagging behind. This approach doesn't work anymore because today almost any candidate can quickly deliver a working solution using Copilot, Cursor, or Claude.
AI-assisted development is already the norm. According to Stack Overflow's 2025 survey, 78% of developers now leverage different AI coding tools. As a result, a candidate's code alone is a weak criterion for measuring their engineering ability.
Many teams that keep using traditional coding assessments face the same pattern: candidates pass vetting, get to work, and ship code but later can't debug or even explain it. This happens because they use code generated by AI without even trying to understand it.
However, the problem isn't in talent but in evaluation. Read on to learn how to vet developers in the AI era and set up interviews that reveal candidates' real performance.
Motivated and focused experts for up to 60% less than locals, delivered in days, not months
Why traditional technical vetting doesn't work anymore
The three most popular technical evaluation techniques have been used for years: take-home tasks, algorithm screens, and live coding. They used to assess real skills.
A take-home task showed whether someone could complete a small project. An algorithm screen evaluated a problem decomposition under time pressure. As for live coding, it checks whether a person can solve a problem while being watched and timed.
All of that worked really well when code writing was the key challenge, and there was no help except documentation or memory. All that has changed with the emergence of AI.
AI agents can now create clean and well-structured code for well-defined problems that are often used in interviews. As a result, it's difficult to assess an engineer's competence if traditional technical vetting is applied.
Coding skills are still valuable, but code itself has become a weak evaluation criterion, as it no longer reflects candidates' critical thinking and judgment.
The danger is that traditional assessments fail slowly, and teams cannot quickly recognize that something is going wrong. They keep ranking and hiring candidates, but more and more of them underperform even though they passed technical checks. If you see this tendency, it's time to change your approach to vetting.
AI impact on traditional technical assessments
Evaluation technique
What it evaluated
AI effect
Ability to independently develop a small-scoped project
You can't really judge a candidate's coding skills, since the solution may fully or partially come from AI agents
Data structure knowledge, problem decomposition, pattern recognition
AI can solve most standard tasks, which is why you cannot be sure whether the candidate can solve them by themselves
Real-time coding fluency, syntax knowledge, design thinking
AI assistance makes code production not a reliable differentiator
So are coding tests still useful? Not completely. They can still show basic technical knowledge, but they cannot isolate the ability we care about most — engineering judgment. That's what a vetting process needs to adapt to.
Skills that have become really valuable
Now that the first draft is usually provided by AI, an engineer's value lies not in producing code but in everything that surrounds that draft — deciding what is good, what is bad, and what holds up in the real world.
That change makes the following skills particularly important:
Problem structuring and scoping: It's still within the engineer's responsibility to decide what to build and how. An AI tool can quickly generate a solution, but an engineer who understands business goals, constraints, opportunity costs, and all the nuances decides if it actually solves the problem.
System design: An engineer is the one who figures out if the components provided by AI fit together in a real product. So, system-level thinking is a must.
Unfamiliar code debugging: Modern codebases usually consist of code written by different engineers, as well as AI-generated code. The ability to analyze, understand, and challenge systems created by someone else is essential to day-to-day engineering work.
AI output review and correction: Evaluating AI-generated code, engineers should be able to spot plausible-but-wrong solutions. This is the new core skill. In fact, it's harder than writing from scratch, because wrong AI code looks right.
Ability to communicate trade-offs to non-techy stakeholders: Many decisions on the project are made by business people. That's why engineers should be able to explain to them why some options are more effective or risky than others. Understanding trade-offs clearly, business stakeholders can make wise decisions.
Taste: This is the developed judgment that helps engineers choose among multiple working solutions the one that is safer, more scalable, and more maintainable.
The core competency is the ability to direct and review AI-generated work quickly without letting quality degrade. The delivery data backs this up. Google's 2024 State of DevOps report states that a 25% increase in AI adoption was associated with a 7.2% drop in delivery stability. Teams generated code faster but shipped more breakage because of the lack of judgment and review.
How to vet a senior developer without being a technical expert
How to assess a developer's skills these days
Whether someone can write code isn't the question anymore. What is important is whether they can understand, debug, and maintain it. This necessitates a technical interview process redesign.
The five techniques described below are of great help.
Evaluating AI-generated code is an activity that can tell you a lot. Give your candidate a 60-100 line pull request that contains some logic bugs, a real security issue, and a defensible-but-wrong design decision. Ask them to go over it. Good candidates catch issues and point out risks. Weak ones spot superficial bugs, such as naming or style problems, but miss critical problems.
This task mimics real engineering work. Present a partially broken codebase and ask the engineer to debug it. Strong candidates always carefully analyze the code, formulate hypotheses, and check them before making any changes.
This technique is a traditional one, but it hasn't become outdated. On the contrary, today, it's even more helpful. While AI generates individual components, the engineer's judgment is needed to design a system that works flawlessly in real-world conditions. A discussion of this kind reveals senior-level engineering skills, letting you understand how a candidate handles trade-offs, scalability, reliability, maintainability, etc.
Ask a candidate if they've ever reviewed their technical decision and why. The answer can show the ability to look back on past mistakes and turn them into valuable lessons. Engineers who have never changed their decision either haven't shipped much or are too self-confident.
Give an engineer a task and let them complete it in real time, using a preferred AI assistant. Your goal is to make sure that they carefully read and examine AI responses, catching mistakes and weaknesses.
Really working evaluation techniques
Technique
Evaluated skills
Execution
Positive signal
Negative signal
Ability to detect issues and evaluate code quality
Give a piece of code with a few bugs, a security issue, and a questionable but defensible design choice
Spotting bugs and risks, prioritizing accuracy, explaining decisions
Identifying style/naming issues, while missing more serious flaws
Debugging unfamiliar code
Ability to spot issues, make improvements, solve problems
Provide a PR with reproducible issues
Hypothesis formulation, systematic verification before code modification
Making spontaneous changes without a clear approach, guessing without validation
Data modeling, critical thinking, senior-level judgment
Discuss a system design problem with real-world limitations and failures
Clear trade-offs, attention to constraints, optimal architecture choices
Vague diagrams, overlooked constraints, over-complicated design
Self-reflection and learning
Inquire about the last time an engineer reconsidered their technical decision and why
Honest reflection, clear explanation, detailed answers
No examples, rigid thinking, inability to justify past decisions
Working on a task with AI allowed
Ability to use AI smartly, ensuring decent quality control
Ask to accomplish a real task with AI aid
Using AI critically, validating responses, catching mistakes, making final decisions on their own
Accepting AI results with little verification, no result ownership
Should candidates use AI in interviews?
This is a hot topic that deserves a separate article, and we actually have one: Should You Let Candidates Use AI in Technical Interviews? Take your time to read it.
The long story short — yes. In most cases, allowing candidates to use AI during technical interviews is totally ok because that's exactly how developers work on a daily basis. If you'd like to evaluate real-world performance, it's thoughtless to restrict AI usage. Instead, allow it but change metrics.
A common mistake is vetting engineers through AI coding while assessing the result as if the candidate wrote every line themselves.
A better strategy is to check how they use AI — whether they verify output, detect mistakes, and take ownership.
By default, an AI-allowed interview should be the norm. Yet, in some situations, AI-free assessments can still be useful. Thus, for certain tasks, such as low-level systems engineering, security-critical work, or compiler development, you may need to make sure executors have fundamental knowledge. The important thing is to understand exactly what signal you'd like to capture.
Avoid operational and HR hassle and the constant threat of attrition with your own R+D department from Devico
A commonly overlooked risk of AI over-reliance
Today, companies are encouraged to prioritize AI proficiency, and that's really important. However, the risk of AI over-reliance is rarely mentioned. The thing is that some engineers ship AI output without even understanding it.
They prompt fluently, demonstrate confidently, and generate expensive, invisible technical debt. They can produce a working solution but can't tell you why it works, can't debug it when it breaks, and can't recognize when the model hands them something wrong.
When used this way, AI can be hazardous. Yes, it has leveled up, but it isn't perfect yet. NYU researchers ran GitHub Copilot through 89 security-relevant scenarios, and about 40% of the offered solutions had security vulnerabilities. And according to Stack Overflow's 2025 data, the biggest developer disappointment cited by 66% of respondents is AI output that is "almost right, but not quite." This is particularly dangerous because clearly broken code usually gets caught, but plausibly wrong code often gets shipped. So, human review is still needed.
When hiring, AI over-reliance is easy to miss, as it looks like real competence at a demo. Both the thoughtful engineer and the over-reliant one deliver working code. The gap shows up later with the first production incident, the first refactor, or the first question about a reason behind a particular design choice.
To distinguish between those two candidates, don't be satisfied with working code. Ask about details, nuances, possible risks, improvements, challenges, etc.
Here's how to check for AI over-reliance:
Ask an engineer to explain AI-generated code, clarifying every detail and its purpose.
Introduce a requirement that AI wouldn't handle, let's say, an unusual input, and check whether the engineer recognizes it.
Take the tool away and ask the engineer to explain the fundamentals.
Ask what parts of the code they would review most carefully before shipping it to production.
Opt for candidates who treat AI-generated code as a draft to verify thoroughly, not a ready-to-deploy solution.
Green flags and red flags during interviews
In AI-assisted development, the most valuable signal during interviews isn't how quickly someone delivers a working solution but how they can grasp the strengths and weaknesses of an existing solution. The signs below can help you separate capable engineers from weak ones.
Green flags:
Verifying every AI response
Understanding code and explaining what should be improved/fixed
Scoping and clarifying the problem prior to coding
Thinking about trade-offs before offering a solution
Red flags:
Using AI-generated code without even understanding it
Inability to explain the presented solution
Failing to debug code when something breaks
Prioritizing a working solution over understanding the problem itself
Assuming that if the code runs, it's 100% correct
Staff augmentation for AI projects: what's different and what to watch for
How vetting by seniority should be changed
AI has changed what real signals look like at each professional level. Vetting should consider this and reflect where real responsibility actually sits at each level and what can really be offloaded to tools versus what must still be owned by the engineer.
Junior level: AI has mostly removed the value of polished syntax and clean boilerplate as hiring signals for juniors, as a model can easily produce them. Instead, you'd better check for learning ability, the habit of verifying rather than trusting, and knowledge of fundamentals that still should be underneath the tooling.
Mid level: Here, the key change is moving from execution to judgment. Assess how engineers work when requirements are vague, how they debug code written by someone else, and whether they can recognize irrelevant or incorrect AI suggestions. A mid-level engineer who accepts every AI response is a junior with better autocomplete.
Senior level / lead: At the senior level, the signal moves even further away from writing code. The main skills are architectural thinking, prioritization, making technical decisions with a lack of information, and quality ownership. One more criterion for senior engineers is the ability to define the norms of AI usage in the team as part of the engineering process.
Changes in vetting by seniority
Level
What to stop testing
What to test instead
Clean boilerplate code, memorized syntax, tidy CRUD apps
Learning speed, clarity of thinking, habit of verifying, knowledge of fundamentals
Ability to produce a working feature
Judgment under ambiguity, debugging unfamiliar systems, AI-generated code evaluation
Isolated coding output, framework expertise, or feature delivery in a vacuum
System design, judgment, decision-making, production ownership, setting team norms for AI usage
Building a standardized process yourself — or working with a vendor who already has it
If you hire engineers quite often, it's worth setting up an internal standardized process with a code review, debugging round, system design, and judgment probe.
The same structure, requirements, and tasks make signals comparable across applicants and interviewers. With a well-thought-out framework in place, you rely less on gut instinct and more on pre-agreed evaluation criteria.
In case you add a couple of new engineers per year or simply don't have the bandwidth to run interviews internally, partnering with an outsourcing company is a good alternative. Yet, it's important to make sure they know how to hire engineers in the AI era and have a reliable vetting process in place. The main question here now isn't whether they test coding skills but whether they vet for judgment and ownership of AI-assisted output.
If the vendor talks mainly about a take-home and a LeetCode screen, they may still hire in the pre-AI way. If their process is similar to one that we outlined above, that's a sign they're in line with how engineering works today.
Takeaway
AI has changed how engineers should be evaluated, making code production a weak criterion and reducing the value of traditional take-home tests. Today, they say little about real engineering competency. What's worth your attention is problem framing, architectural thinking, unfamiliar code debugging, and other skills needed for proper AI output verification.
Teams that don't adjust their evaluation approach risk hiring people who pass the test but can't do the job, over-relying on AI.
Devico's vetting already reflects the new reality: engineers are evaluated on AI fluency as well as judgment and ownership of AI-assisted work.
If you're looking for a partner who understands how to vet developers in the AI era and organizes their process accordingly, we're happy to support you with our staff augmentation services.
Great software starts with great people