The model came back in 36 hours. Three statements, a DCF, a sensitivity table, a two-page memo with footnotes. Cleaner than the one your second-year associate built last month. The candidate is a junior from a school you do not usually recruit from, and the phone screen was fine but not this fine.
You have two choices. Advance someone who may not be able to do the work, or reject someone who may be the best candidate in the pool. Either way you are guessing, and if the guess is wrong, the hiring manager will not remember that the deliverable looked perfect. They will remember who advanced the candidate.
This post is about how to stop guessing.
What the submission itself can tell you
Reviewers have developed instincts for AI-built work, and some of them are sound. In a model, look for formulas that are technically correct but structurally odd: hard-coded values where an analyst would link cells, consistent formatting across sections that a human would have built at different times, assumptions that are reasonable in isolation and never reconciled to each other. In a memo, look for prose that is even in tone from the first paragraph to the last, confident transitions, a conclusion that restates the question, and no sentence that sounds like a person who has run out of time.
Then look at what is missing. Human analysts leave traces of the order they worked in: a scratch tab, a note to self, a version-two label, a number that was updated in one place and not another. AI-generated work tends to arrive finished. A deliverable with no history is not proof of anything, but it is worth noticing.
Statistical classifiers can do a version of this at scale. Tools that measure perplexity and burstiness in prose, or naming entropy and comment cadence in code, pick up the fingerprint of machine-generated text with reasonable accuracy on long documents. They are weaker on short answers, on heavily edited output, and on candidates who write in a second language, and every one of them produces false positives.
Which is the point. The submission can raise a flag. It cannot, on its own, tell you who did the work.
Why reviewing the deliverable is not enough
Three things break the review-the-file approach.
The first is that a candidate who uses AI well does not submit raw output. They prompt, read, edit, re-prompt, and paste pieces into their own structure. The result has human fingerprints on it because a human touched every part of it. Detection on the artifact gets you to "probably some AI involvement," which is also true of most work your current employees produce, and your policy almost certainly allows some of it.
The second is that the deliverable tells you nothing about who was in the room. The candidate's friend who works at a fund, a paid service, a tutor on a remote-desktop session: none of them leave a statistical signature in the spreadsheet. The same remote-access tools that show up in live-interview fraud (we covered the live-interview version in How to detect if a candidate is using AI in a remote interview) are used on take-homes, with more time and less scrutiny.
The third is that an accusation based on the file alone does not hold up. If you remove a candidate because a classifier scored their memo at 78 percent likely AI, you have a decision with no defensible record, and if that candidate pushes back, you have a problem your legal team will not thank you for.
So the file is one layer. You need the other two.
How to actually tell who did the work
The firms that have solved this run three controls together, and the order matters.
Set the rule and say it will be checked. The case study instructions state what is allowed (research, reading, spell-check) and what is not (generating the analysis, the model, the memo or the slides; anyone else touching the work). They state that submissions are reviewed for AI-generated content and that the candidate may be asked to walk through any part of the work live. Most of the deterrent effect comes from this paragraph. It also means a candidate later removed from the process agreed to the terms.
Record what ran on the machine during the window. This is the layer most firms do not have. A lightweight desktop agent runs on the candidate's laptop for the duration of the take-home. It does not read files, capture content or record the screen. It records whether restricted programs ran: AI assistants, overlay tools such as Cluely and Interview Coder, remote-access software such as TeamViewer and AnyDesk, virtual machines. The finding is binary and comes with the process name, hash and timestamp. If a restricted program ran, you know and can show it. If none did, the candidate has a clean record they can point to, which the good candidates like.
This is what ScreenComply does for financial services recruiting teams: the agent during the window, AI-generated-content review on the submission, and an integrity report per candidate that goes into the file alongside the work. The reviewer sees the outliers first, ranked against the cohort, instead of reading three hundred models cold. Details on the finance deployment are on the financial services page.
Verify the author live. Fifteen minutes on video within a few days of submission, with live detection running. Pick one cell, one assumption, one slide. Ask the candidate to change it and explain what moves. Someone who built the model does this without thinking. Someone who did not will stall, and you will have watched it happen rather than inferred it from a score.
Together the three produce something the file alone never could: a record. Policy agreed, machine clean or not clean, author verified or not verified. The hiring committee decides, and the decision has evidence under it.
Keep the take-home
Most advice on the web right now says to abandon take-home assignments because AI made them meaningless. For finance, consulting and professional-services recruiting that advice is wrong. The take-home is where you see how someone structures a problem over hours rather than minutes, and it is the only stage that scales across a campus class of several hundred. Dropping it throws away the signal because the proof of authorship broke. Fix the proof.
The full process, with a checklist, is in How to prevent AI cheating on take-home case studies and modeling tests. Candidates, in our experience, prefer a monitored take-home to losing the stage entirely to a timed live test: it keeps the part of the process where they can show their best work.
Frequently asked questions
Can software detect if a take-home was written by AI?
Partly. Classifiers on the submitted document, spreadsheet, slides or code find statistical signs of machine-generated work and report them as a ranked signal with evidence, not a verdict. They are strongest on long documents and weakest on short or heavily edited ones, and they produce false positives. Combined with a record of what ran on the candidate's machine during the window and a short live follow-up, the firm can establish authorship with a defensible record.
Does this work for Excel modeling tests?
Yes. The desktop agent runs on the candidate's machine regardless of the application, so an Excel model is covered the same way as a Word memo or a PowerPoint deck, and the spreadsheet itself is reviewed on submission.
What does the candidate have to install?
A desktop agent of about 100 MB on macOS or Windows, for the duration of the take-home window, removed afterward. The candidate is told what it monitors before the window opens. It captures system-level signals only: no files, no keystroke content, no continuous screen recording.
Is monitoring a take-home legal?
Candidates are notified and consent before the window opens, the system records system-level signals rather than personal content, and retention is configurable, including analyze-and-discard. The firm's legal and compliance teams configure the deployment and should review the candidate notice for their jurisdictions. ScreenComply is not a law firm.
Should we replace take-homes with live tests instead?
Not for roles where hours-long structured analysis is the job. Live tests measure speed under observation; take-homes measure depth. Keep the take-home and make authorship verifiable.
ScreenComply provides interview, exam and take-home integrity software for financial services recruiting teams, universities and assessment platforms. See how a monitored take-home runs: Book a demo.
