
A polished submission is an incomplete hiring signal
A candidate can hand over a fluent AI-assisted memo without revealing whether they understand its support. Another can identify a serious gap but present the correction awkwardly. If the panel rewards appearance alone, it may miss the very behavior it wanted to inspect.
My starting point is a narrow work sample with visible sources and a defined recipient. Ask the candidate to review an output, correct what the record supports and explain what remains unresolved. Then assess those actions separately. This makes the conversation more specific than asking whether someone is good with AI.
The scorecard below is an AI Vortex editorial proposal demonstrated with a synthetic exercise. It has not been validated against job performance, tested for selection outcomes or approved as a hiring instrument. Its first use should be a panel rehearsal that reveals whether the task and criteria make sense for the position.
Test the work you need on entry
The U.S. Office of Personnel Management's work-sample guidance describes tasks resembling job activities and cautions that this method may not fit competencies an employer intends to teach after selection. That distinction matters for a junior role: difficulty with an unfamiliar proprietary interface can tell you little about the source judgment you planned to assess.
Write the entry expectation before designing the exercise. For a research-support position, it might be preserving the link between a fact and its source and escalating an unresolved conflict. For a role requiring substantive legal analysis, it will be different and needs a qualified subject-matter reviewer. A generic legal AI test cannot stand in for both.
Give the panel a practical constraint: if it cannot explain why an observed behavior is necessary for this role, leave that behavior out of the score. Tool speed, confident delivery and elaborate formatting should not enter the assessment by accident.
An exercise with enough evidence to inspect
Here is an independent synthetic rehearsal. Record A is an email stating that a meeting occurred on May 6. Record B is a letter stating that the same meeting occurred on May 5. Record C is an undated attachment. The generated summary says: ‘The meeting and attachment are dated May 6.’ No source establishes the attachment date.
The instruction is to produce a short corrected note for a supervising attorney using only those records. The note must identify support, preserve unresolved questions and name a next step. This task asks for record review; it does not ask the candidate to resolve a legal issue or operate a vendor product.
A defensible rehearsal answer would say: ‘Records A and B conflict on the meeting date: A states May 6 and B states May 5. Record C is undated. I have not assigned a date to the attachment or resolved the meeting date. Please confirm whether another record can establish either point before the chronology is relied on.’ This answer preserves what the evidence permits without turning uncertainty into a conclusion.
The synthetic review lab provides additional synthetic practice scenarios and explanations. Because its answers are public, do not reuse them unchanged as a confidential candidate test. Use the lab to train reviewers or discuss method; a real assessment needs a fresh, reviewed task matched to the role.
Score dimensions, not an overall impression
Use 0 for not demonstrated, 1 for partly demonstrated and 2 for demonstrated in this exercise. These anchors are editorial design choices, not calibrated measurements. Record the supporting sentence or action for each rating. Use “not observed” when the task did not elicit a behavior; do not silently convert missing evidence into failure.
| Dimension | 0 — not demonstrated | 1 — partly demonstrated | 2 — demonstrated |
|---|---|---|---|
| Source fidelity | Repeats an unsupported date as fact | Notices a date problem but does not connect it to the records | Links both meeting dates to A/B and keeps C undated |
| Uncertainty handling | Chooses one date without support | Expresses doubt but leaves the conclusion ambiguous | Preserves the conflict and distinguishes it from missing information |
| Correction | Leaves the original error in the deliverable | Removes one error while leaving another material one | Revises the meeting statement to preserve both attributed dates and explicitly leaves the attachment undated |
| Handoff | Offers a clean-looking conclusion without a review request | Requests review without identifying the open point | Names the unresolved point and a relevant next inquiry |
| Explanation | Cannot explain the support for the correction | Explains part of the change but misses its limitation | Explains why the records do not permit either unsupported conclusion |
What the ratings reveal in the synthetic example
The rehearsal answer above demonstrates source fidelity, uncertainty handling, correction and handoff under these anchors. Its explanation rating remains not observed until the reviewer asks why the correction was necessary. That distinction prevents the panel from awarding evidence it has not collected.
Now compare a shorter response: ‘I corrected the date to May 5 and left the attachment undated.’ It fixes the unsupported attachment date but resolves the meeting conflict without a source-based reason. The panel can recognize a partial correction while examining the unsupported choice. A single polished-versus-unpolished judgment would hide both observations.
Do not add the ratings into an automatic pass mark. A total can allow good formatting or communication to obscure a material unsupported claim. Use the separate observations to decide what follow-up is needed, while keeping the final selection within the employer's appropriately designed process. These example ratings show how the proposed anchors work; they do not establish predictive validity.
Make the follow-up comparable
OPM's structured-interview guidance uses predetermined questions and common rating standards. For this rehearsal, prepare the follow-up in advance: which record supports each change, what remains unknown, and what additional information would permit a stronger conclusion? Ask the same core questions of each participant.
Set the task conditions before use: source packet, permitted tools, deliverable, time expectations and how clarification questions are handled. Provide the environment needed for the task instead of assuming candidates own a paid subscription. Record assistance or technical interruptions so the panel can distinguish task performance from conditions it created. Accommodations may require changes to administration; consistent assessment does not mean ignoring those needs.
Have two reviewers rate the rehearsal independently and compare the evidence behind disagreements. If one treats a cautious answer as failure and another treats it as sound escalation, fix the ambiguous anchor before involving candidates. Agreement in this rehearsal tests whether the panel understands the instructions; it does not validate the hiring procedure.
The employer still owns the selection method
The EEOC technical-assistance page on employment tests discusses job relevance, validation, discriminatory effects and reasonable accommodation. It also makes clear that employers retain responsibility for the tests they use. A public template does not settle those obligations, and the page is U.S. guidance rather than a complete rulebook for every jurisdiction.
Before using a work sample in actual selection, involve the people responsible for job analysis, assessment and applicable employment requirements. Confirm that the task tests an entry expectation, can be administered appropriately and has evidence supporting the intended use. Do not collect client material or ask candidates to produce unpaid work for a live matter.
I would judge this scorecard first by what it makes the panel notice. Can reviewers distinguish unsupported confidence from a properly bounded answer? Can they identify the exact evidence they still need? If not, the hiring team has more design work to do before it asks candidates to perform.
For the candidate-facing delivery method, see AI skills for new associates. For the labor-market question, see associate hiring evidence. A work sample answers a narrower question about observed task performance; it cannot explain national hiring trends.
Questions and answers
Is this a validated hiring assessment?
No. It is an editorial scorecard and synthetic demonstration. Employers need a job-relevant, appropriately supported selection process before using any exercise to make hiring decisions.
Should we require candidates to use a particular AI product?
Only if that product skill is genuinely part of the entry requirement and the assessment is appropriate. Provide the required access; do not assume a candidate has a paid account.
Can the review lab be used as a hidden-answer candidate test?
Its explanations are public, so an unchanged exercise may test familiarity with the answer. Use it for reviewer practice or transparent discussion; develop a fresh, reviewed task for selection.
What does a not-observed rating mean?
The exercise or follow-up did not supply evidence for that dimension. It is different from observing a failure. Ask a relevant follow-up rather than manufacturing a score.
Sources and scope
- U.S. OPM: Work Samples and Simulations. living first-party guidance; checked 2026-09-08.
- U.S. OPM: Structured Interviews. living first-party guidance; checked 2026-09-08.
- EEOC: Employment Tests and Selection Procedures. 2007-12-01 technical assistance; current hosted page checked; checked 2026-09-08.
U.S. occupational and professional sources inform this guide. Local rules, qualifications and employer requirements differ. Examples and practice plans are editorial proposals; they are not employment forecasts.
Editorial update. New employer-facing editorial rubric with a worked synthetic example. OPM and EEOC inform the assessment boundaries; they do not endorse or validate this scorecard.
Prepared with AI-assisted research and editorial verification for AI Vortex. Sources are linked where claims are made.