What to Do When AI Can Solve Your Take-Home

Sep 14, 2026 · 5 min read

MB SamuelFounder
What to Do When AI Can Solve Your Take-Home

TL;DR: When candidates use AI to complete take-homes, it's harder to differentiate the results. The models and slides all look credible, but it tends to fall apart during the live review. To differentiate, consider an AI-assisted take-home, where you look at the AI session and prompts instead of only the deliverable. This helps you understand the process and the thinking behind the output.

At Gradient, we often hear that take-home assignments teams have used for years for roles like finance, marketing, sales, or biz ops, are no longer working.

Excel analyses that used to differentiate candidates are now easy for Claude. Codex creates polished slides that can mask limited understanding. Teams used to worry about calibration and reviewer time. Now they usually start with some version of: we think candidates are running this through AI, and we're not sure what to do about it.

Test your own take-home first

Open the take-home you send candidates today. Paste the brief into Claude or ChatGPT with the same attachments candidates receive, and read what comes back.

The output is rarely excellent, but it's usually strong enough that you need a live follow-up session, which means scheduling and setting aside team time, just to understand if a candidate applied their own thinking or relied blindly on AI outputs. The assignment keeps producing scores, and those scores now reflect some mix of candidate skill and model quality that you can't pull apart by reading the submission.

Why AI detection tools don't work

Often, teams respond to this by looking for a detector. Unfortunately, tools that claim to identify AI-written text are unreliable, and research on them has repeatedly found they misflag writing by non-native English speakers, so false positives land hardest on candidates who already face more friction in your process. These also don't work well for formats like Excel.

Knowing a model was involved tells you very little. Knowing how it was used tells you a lot.

What to look at instead of the final deliverable

Picture two candidates submitting similar memos.

The first wrote three sentences of prompt, took the draft, and changed the greeting. The second gave the model the account history, asked it to argue the opposite recommendation, found a figure that didn't hold up, checked it against the source, and rewrote the middle section by hand.

Graded on the memos alone, they score about the same. But when you ask follow up questions, probe into the decisions, or extend the work further, one candidate will vastly outshine the other.

Three ways teams are rewriting the take-home

Give every candidate the same AI

Most take-homes today either ban AI, which is hard to verify, or don't set guidelines, in which case the candidate with the more expensive subscription tends to submit the better work.

Providing the environment yourself gives everyone the same model, the same tools, and the same time limit, so one rubric is comparing work produced under the same conditions. We went deeper on this in Not every candidate has the same AI.

Look at the session, not only the submission

Prompts, edits, and corrections are where judgment is visible. You can see whether a candidate verified a claim, whether they pushed back on a confident wrong answer, and whether they knew when to stop revising.

Write tasks AI can't finish alone

The briefs that hold up in 2026 share a few qualities: something the candidate needs is missing and they have to notice, two people in the scenario want different outcomes, a figure in the source material is wrong, and the final call is defensible in more than one direction.

AI will answer all of those confidently, and often incorrectly. The candidate has to catch it, and catching it is the behavior you're trying to hire.

Try it

Two candidates, one brief: a messy customer inbox, and a request to recommend what the team should build next. What makes the second prompt better?

Candidate A

Read these support tickets and tell me the top 3 feature requests.

Candidate B

Six questions to audit your current take-home

You don't need new tooling to answer these.

  1. If you paste the full brief into a frontier model, does the output clear your bar?
  2. Does the brief state, in writing, whether AI is allowed?
  3. If AI is allowed, does every candidate have access to the same tools?
  4. Does anything in the brief require a candidate to catch missing or incorrect information?
  5. Can two reviewers score the same submission and land in the same place?
  6. If a candidate did excellent work, could you tell from the artifact alone?

Where to start

Most teams start by requiring candidates to complete their take-homes live on Zoom, or using follow-ups to ask detailed questions. Both of these are effective, but they use a lot of your team's time, and they require scheduling live sessions and sometimes paying for flights.

Gradient standardizes the model and harness for each assessment so that the environment is fair and consistent. Your reviewers get the deliverable plus the full session behind it: every prompt, edit, and correction. This lets you see the thinking - the inputs - as well as the final verdict - the output. It's async, and quick to review.

If you're rethinking your take home and want to chat about a better option, please reach out to learn more.

Start hiring AI fluent talent.

We’ll show you an assessment, walk through the scoring engine, and get you live in a few hours.