Back to all articles
Hiring guides

AI Fluency Assessment: How to Test AI Skills in Hiring

An AI fluency assessment shows how a candidate actually works with AI, from framing the task to catching its mistakes, which no resume line can tell you.

H
The Hyrr team
•8 min read

A resume line that reads "proficient in ChatGPT" tells you the candidate has used the tool, or wants you to think so. It does not tell you whether they get good work out of it, notice when it is confidently wrong, or know when to leave it closed. A self-rating tells you even less.

An AI fluency assessment closes that gap. You watch how someone actually works with AI on a realistic task, and you score what you see against anchors written in advance.

Why AI fluency has to be tested, not asked about

For many roles, AI tools are now part of the daily work: drafting, coding, summarizing, analyzing. How well someone uses them is a real job skill.

Picture two candidates with identical resumes. One pastes the task into a chat window and submits whatever comes back. The other explains the goal and audience, splits the job into steps, checks the numbers against the source, and catches an invented citation before it reaches a client. On paper they look the same. At work, one of them creates rework and risk.

AI literacy is knowing what the tools are and what they do well and badly. AI fluency is using them well, under realistic conditions, to produce work you would put your name to. An AI literacy test can check the first. Only an exercise can show the second.

What AI fluency looks like: seven observable behaviors

What you cannot observe, you cannot score consistently, so define AI fluency as behaviors you can see or hear:

  • Framing the task. They state the goal, audience, constraints and format, and supply the context the tool cannot know: the data, the house style, the edge cases.
  • Breaking the work into steps. They split a large task into pieces the tool handles well, instead of asking for everything at once.
  • Iterating. They treat the first output as a draft, give specific corrections, and change approach when one is not working.
  • Verifying. They check facts, figures, code and sources against something real: the original data, a test, the documentation.
  • Knowing when not to use AI. They recognize tasks where the tool is slower, riskier or the wrong instrument, and do those themselves.
  • Handling data responsibly. They keep confidential and personal information out of tools not approved for it, and follow the employer's policy.
  • Owning the result. They can explain and defend every part of what they submit, whichever parts the tool drafted.

Why the usual approaches fall short

Banning AI in the interview. A blanket ban tests a way of working the job may no longer use, like assessing an analyst without a spreadsheet, and tells you nothing about how the candidate uses AI. If one part must be done unaided, such as reading code, make that a deliberate choice and say so up front. For keeping those parts honest, see our guide to preventing cheating in remote interviews.

Trivia quizzes about tools. Questions about model names, settings or features test a fast-moving vocabulary. A candidate can ace them and still paste a client's contract into a public chatbot. A short AI literacy test has a place in training; as an AI skills assessment for hiring, it measures recall, not practice.

Judging only the final output, or the speed. AI makes polished output cheap. Two identical deliverables can hide very different processes: one checked line by line, one lucky. Rewarding speed favors the candidate who skipped the checks. Score the process and the reasoning, not just the result.

How to assess AI skills in candidates: three methods that work

1. AI-allowed work samples, with the process in view

Give the candidate a small, realistic task from the role and explicitly allow AI tools. Make the process visible with a shared screen, or ask them to walk you through their prompt history afterward. Score how they framed the task, what they changed and what they checked, not just the final file. Keep the task small enough to finish comfortably, so you see deliberate work rather than a race.

2. "Find the flaw" exercises

Hand the candidate an AI-produced draft with planted errors, and ask them to review it as if it were going out under their name. It might be an analysis with a wrong total, code with a subtle bug, a summary that misstates its source, or a job ad with an invented requirement. Vary the difficulty of the errors, keep an answer key, and use the same draft for every candidate. It targets verification directly and usually takes less time than a full work sample.

3. Follow-up questions about their choices

When testing AI skills in interviews, the most revealing part is often the conversation after the exercise. Ask:

  • "Why did you trust that output?"
  • "What did you check, and how?"
  • "Where did you decide not to use the tool, and why?"
  • "Tell me about a time an AI tool gave you an answer that looked right and wasn't. How did you find out?"

Strong candidates give specific answers you can check against what you watched. Weaker ones describe what the tool did, not what they did.

An AI fluency assessment rubric you can adapt

A rubric turns impressions into evidence. Write the anchors before you see the first candidate, and use the same ones for everyone. Structured interviews predict job performance about twice as well as unstructured ones (Sackett et al., 2022), and written scoring anchors are a core part of that structure. For how to write anchors step by step, see our structured interview rubric template.

DimensionNoviceCompetentProficientExpert
Task framingPastes the task as given, with no contextStates the goal and format; adds context when askedGives goal, audience, constraints and sources up front; splits the work into stepsPlans the approach first and decides which steps suit AI at all
IterationAccepts the first output, or retries at randomRefines with vague requests such as "make it better"Gives specific corrections; changes approach when stuckWorks toward a clear standard and stops when the work meets it
VerificationDoes not check; misses the planted errorsCatches surface errors such as typos and obvious slipsChecks figures, sources and code against data or tests; catches most planted errorsChecks systematically, catches subtle errors, and explains how each was found
JudgmentUses AI for everything, or refuses it without a reasonAvoids AI for clearly unsuitable tasks when promptedChooses where AI helps and where it does not, and explains whyWeighs speed, risk and quality at each step; justifies any choice not to use AI
Responsible usePastes sensitive data into tools without a second thoughtAvoids obviously sensitive data when remindedMasks personal and confidential data unprompted; follows the stated policyFlags data and policy risks in the task itself, and owns every line of the result

A few rules make the rubric work in practice:

  • Score each dimension separately, from what you saw or heard, and write the evidence next to the score.
  • Weight dimensions to the role. Verification matters most for an analyst; responsible use weighs more for anyone handling candidate or customer data.
  • Set a floor for critical behaviors. A candidate who pastes personal data into an unapproved tool should not pass because their framing was excellent.

Example exercises by role

Software engineer. Give a small repository with a bug to fix, and allow an AI coding assistant. Watch whether they read generated code before running it, run the tests, notice a call to a function that does not exist, and can explain every change. Ask which suggestion they rejected, and why.

Data analyst. Provide a small dataset and a business question, with AI allowed. Plant a trap: duplicated rows, mixed currencies, or a misleading column name. Score whether they sanity-check results against the raw data, catch the trap, and state their confidence honestly. The key follow-up: "How do you know this number is right?"

Recruiter or marketer. Ask a recruiter to turn a hiring manager's notes into a job ad with AI's help, then review the draft for inflated requirements, exclusionary wording and promises the company cannot keep. A marketer can do the same from product notes, checking for invented features and statistics. Either way, ask what they would never paste into a tool.

People manager. Give a scenario, such as turning anonymized team feedback into a development plan, with AI allowed. Watch their judgment: what they hand to the tool (structuring notes, drafting an agenda) and what they keep for themselves (a performance verdict, a sensitive conversation). Then ask how they would set expectations for AI use on their team.

Keep the assessment fair

An AI fluency assessment is a selection step like any other, and it needs the same care.

  • Same tools, same time, same instructions. If one candidate has a paid assistant and another a free one, you are partly testing their subscriptions. Provide the tool, or state exactly what is allowed.
  • Tell candidates in advance what is allowed, for which parts, and what you will look at. A surprise measures nerves, not skill.
  • Do not penalize a justified decision not to use AI. A candidate who writes a short query by hand because it is faster, and can say why, is showing judgment, which is on the rubric.
  • Score substance, not style. Score the context they gave and the checks they made, not whether their prompts look like yours.
  • Watch outcomes over time. Compare pass rates across groups. Under the four-fifths rule in the Uniform Guidelines on Employee Selection Procedures, used by the US EEOC, a group's selection rate below four-fifths of the highest group's rate is generally regarded as evidence of adverse impact.

How Hyrr approaches AI fluency

Hyrr's Practical AI Fluency assessments are new and rolling out. They are in-interview AI exercises and hands-on labs that measure how well a candidate works with AI tools.

More broadly, Hyrr's interviews already rest on the principles in this guide:

  • Anchors written in advance. Every job has a scorecard of weighted skills, each with a written grading guide (behaviorally anchored rating scales, or BARS) that defines what a Novice, Competent, Proficient and Expert answer looks like, at a depth the team sets using Bloom's taxonomy. The rubric above uses the same four levels.
  • One standard for everyone. The guide is written before the first applicant, and the same guide is used for every candidate.
  • Attention to the work, not only the answer. In hands-on work, the candidate uses a shared code editor or a whiteboard. The AI interviewer sees the code as it is typed and the diagram as it is drawn, and asks about it. The work is saved to the report.

To ask about Practical AI Fluency for your roles while it rolls out, reach the team through the form on the Hyrr home page; they reply within 24 hours. To see the live AI interview from the candidate's side, take the demo interview: pick a sample job and take a real interview in your browser or by phone, in minutes. You can also try the frontend engineer demo or the data analyst demo.

See a live AI interview

Pick a sample role and take the interview yourself, in the browser or by phone. Your scored report lands in your inbox.