AI Fluency and the War Against AI: Some Thoughts of Mine

Product Partner
TL;DR: AI detectors are trying to answer a question that stopped mattering. How well somebody works with AI is the one that didn't. I took an AI fluency assessment twice, 6 months apart, and went from 76 to 89. What makes those numbers mean anything is the method: a real conversation scored against a published rubric, so it reads how you actually work instead of what you claim on a resume. It's still a screen, and the final call belongs to a person watching you defend your own work. One of the criteria also fell 22 points on knowledge I never lost, which turned out to be the most useful thing it told me.
A while back Substack asked me to declare how much AI I use when I write. This is what I put:

I meant it, and it's still the easy version of the argument. Judging my own work is simple. I know what I did, I know what I checked, I know which parts I'd defend in a room.
The hard version arrives when you have to judge somebody else's AI fluency from their work. Fast, at volume, and with no visibility into how any of it was made.
I've been sitting with that problem all year, so I'm writing a few pieces about what I'm actually doing about it: the things I'm trying, the ones that worked, and the ones that didn't. Scaling a services business is hard. Scaling one from Pakistan is harder. Scaling one while every candidate, every application and every deliverable arrives with an invisible AI layer baked into it's a different job from the one I learned.
This first piece is about AI fluency: what it's, why hiring can't read it off a resume anymore, and what happened when I put a number on my own twice, 6 months apart, and watched one of those numbers go the wrong way.
What AI Fluency Actually Is
Strip the jargon off and AI fluency is simple. It's the difference between people who use AI to do better work and people who use AI to produce slop.
Both of them are using the same models. Everyone keeps missing that part. The tools are commodity now, $20 a month, available to your best engineer and your worst one on identical terms. What separates them is judgment: knowing what to hand over, knowing what comes back wrong, knowing when the confident answer on your screen is confidently wrong.
AI multiplies whatever skill you already have. An expert with AI is 1 times 10. A beginner is 0 times 10, and 0 times anything is still 0. I'm not an AI skeptic, the multiplier is real and I use it every day. But a multiplier only acts on what you bring it, which is why the tool was never the interesting variable.
There's real research on this now, which is newer than most people realize. Anthropic published an index earlier this year built on close to 10,000 real conversations, and 2 findings from it are worth your time.
The first is structural. They defined 24 behaviors that mark genuine AI fluency, and only 11 of them are visible inside a conversation at all. The other 13 happen somewhere you can't see: whether the person checked the output against a source, whether they told anyone they used AI, whether they understood the thing they shipped. More than half of what makes somebody good at this is invisible in the artifact.
The second finding should worry anyone who hires. When AI produces a polished output, people get worse at scrutinizing it. They check fewer facts and question the reasoning less. Anthropic doesn't claim to know why. Their guess is that if the work looks finished, people treat it as finished.

Source: Anthropic, The AI Fluency Index, 2026, 9,830 conversations analyzed.
Put those 2 findings together and you have the whole problem. Hiring runs on the artifact. The artifact is now cheap and polished, and it actively suppresses the scrutiny that would catch it.
We Are Fighting the Wrong War
The industry's answer to this has been detection, and detection is a dead end.
In July, Substack shipped an AI detector from Pangram across the platform, so anything over 100 words now gets scored and readers see the verdict: human, AI-assisted, or AI-generated. Writers called it a witch hunt almost immediately. They were right to.
Pangram's CEO puts the false positive rate at about 1 in 10,000. That number can't mean what it sounds like it means. A false positive rate is only meaningful at a stated threshold, and when the output is a sliding score rendered as a 3-way label, the threshold is doing all the work and nobody published it. We've seen this film before. Turnitin claimed a false positive rate under 1% for its AI detector. When the Washington Post tested it in 2023, it got more than half of a 16 sample set at least partly wrong. A tiny sample, which is the point: nobody outside these companies has the data to check the claim.
Then there's who gets caught. Detectors misfire on people writing in a second language, which is most of the people I work with and, on a bad day, me. I run a business in English from Pakistan. The machine deciding whether my writing is real is measurably worse at reading people like me.
But the deepest problem isn't accuracy. The question itself is wrong.
"No AI was used" isn't a category that exists anymore. Nearly everyone I know runs research through a model before they write a word. Over at SK NEXUS (my publication), we use AI to fact check a lot of our own articles, especially the deep ones that can carry 30 to 40 facts in a single piece. Every claim needs a source, and that's slow work by hand. This article was drafted with AI, from an hour of me ranting into a microphone. None of that's a confession. It's just Tuesday.
What people actually object to is slop: careless, unowned output that wastes the reader's time. Slop is a judgment about quality, and quality is exactly the judgment a detector can't make. It can tell you a machine was involved. It can't tell you whether anybody was thinking.
Detection asks whether you used AI. Fluency asks how well. Only one of those questions has an answer worth paying for.

What an AI Fluency Assessment Measures
I've spent most of this year trying to get better at this myself. Less noble than it sounds. My delivery quality is the product I sell, and now there's a team whose delivery quality is the product too. How good we actually are with AI, and where the holes sit, is a standing item on my list rather than a project with an end date.
Early this year a peer I had worked a contract with sent me something he had built for measuring exactly that. It's an AI fluency assessment called AISA, and the interesting part is the format. It's a conversation, 20 to 40 minutes, with an AI interviewer that asks what you actually do and then keeps pulling the thread. No multiple choice, and no survey asking you to rate your own prompting out of 10.
That distinction matters more than it sounds. A paper from this year's Learning Analytics conference, Zhang and colleagues, measured how well self reported AI competence predicts measured competence. The correlation came out between 0.07 and 0.23. Close enough to nothing. Their population was teachers rather than job candidates, so take the size with salt and the direction seriously: asking people how good they're at AI tells you almost nothing about how good they're at AI. Every resume line claiming AI proficiency is a self report.
Under the hood it runs 2 tracks. One AI holds the conversation and never sees your scores. A second scores every message against a published rubric, pulls the exact quote that justified each score, and passes steering notes back so the next question probes wherever the evidence is thin. The candidate only ever meets the friendly one. It's a good piece of design, and it's the part I liked most the first time around.
I took it in March, before I had built any of the agentic workflows I run now.
76 out of 100. Proficient tier. The Builder. That's exactly where their own distribution draws the line for the top 10%, against roughly 1,900 people whose average is 46. (My original report said 73. The dashboard now says 76, because they re-scored old sessions when the methodology changed. Worth knowing that a score isn't a fixed object.) I was pleased with myself for about a minute, which is how long it took to scroll to the dimension breakdown.
Workflow 85. Prompting 75. Technical 75. Critical Thinking 70.
Safety and Responsibility: 45. My weakest by 25 points, and the only one sitting in the Developing band.
I went looking for what I said to earn that. The interviewer had asked what would make me hesitate before sending AI output to a client. My answer, verbatim, typos and all:
well, my eyes? i hav eover 10 year sof experience, its pretty easy to know when the output is wrong
The scoring note was polite and completely correct: awareness that output needs checking, but no systematic guardrails, no protocols, just personal judgment. It's the answer of somebody who has never had to hand the job to anybody else. Which, at that point, I had not.
That's the moment the tool earned my attention. Not because it flattered me. Because it found the one part of my practice I couldn't have defended, and it found it inside a 37 minute conversation.
So I Took It Again
6 months later, in the middle of writing this article, the honest thing to do seemed to be taking it a second time and publishing whatever came back.

A lot changed in between. I moved most of my work into agentic workflows, Claude Code wired into version control, rules and context managed as real files. I went from using AI to running it.
89 out of 100. Advanced tier. The Architect. Top 3%
It asked me to predict my score before revealing it. I guessed 80 and scored 89, which puts me on the side of their population that underestimates itself. Their own data says the least fluent overestimate by around 40 points and the most fluent underestimate by around 27. Make of that what you'll.

The part that convinced me it's measuring something real sits underneath the headline number.

Safety went from 45 to 82. The report quotes the answer that did it, where I described the review gates now built into my workflow, every step pausing for confirmation before it overwrites anything. The scorer called it a designed checkpoint system rather than ad hoc caution. In March my safety practice was my eyeballs. In September it's a system. That's a real change in how I work, and the test caught it.
Now the part a testimonial would leave out.
Critical Thinking moved 2 points. Output Evaluation moved minus 1. The dimensions I didn't work on didn't move. If that 13 point jump were only me knowing the game better the second time, everything would have drifted up together. It didn't. The instrument moved where I moved and sat still where I didn't, and that discrimination is better evidence for it than my score is.
And then this. Tool Landscape fell from 81 to 59. A 22 point drop, on a criterion measuring how well I know the AI tools on the market, across 6 months in which I unambiguously learned more tools, not fewer.

The report explains itself, to its credit: only one tool mentioned, no comparison to alternatives, organic only, needs probing. In March I happened to reel off 6 tools. In September I described one consolidated setup, because that's how I work now. Same knowledge. Different conversation. 22 points.
That's the real limitation of this format, and it cuts both ways. In my first report several criteria were flagged as not yet probed and still scored 81 to 85. One of those inflated scores corrected itself downward the moment a second conversation actually went there. Adaptive questioning is what makes a conversation better than a quiz, and it's also why 2 sessions with the same person produce different numbers. Coverage depends on where the conversation wanders.
I should be straight about my own 13 points too. 3 things are inside that number: I genuinely improved, I had read my own report and their published rubric, and my first session ended early with criteria left unprobed while the second ran to full coverage. The tool can't separate those, and neither can I. Their own quality page concedes that test-retest data is still a planned next step. As far as I can tell this is the closest thing to public test-retest data the product has, and it comes from the least clean subject available: somebody who had studied.
Where the Machine Stops
I have one real criticism, and it's the interesting kind.
One thing to place first. My first run was in March, roughly a month before AISA opened to the public, so I was an early tester on an early build. By the second run their own site was reporting more than 2,500 assessments run. What follows is real and it's still the part I find most interesting, but I hit it in a version almost nobody else saw.
Late in my first session I told the interviewer it had run out of context window and was losing track of the conversation. It replied that it had the full conversation in view and no context issues at all. Then the report scored my line as strong evidence of limitation awareness, crediting me with identifying a specific failure mode as it happened.
It scored the claim. It did not score whether the claim was true. Read that again, because it's the whole problem in one line: the system's own conversational track had already contradicted the claim, and the scoring never reconciled the two.
For what it's worth I still think the context was degrading, and that reading isn't kinder to the machine. Either it credited a false claim, or its own account of its own state was wrong while the scorer took my version anyway. Either way a contradiction sat in the transcript and nothing resolved it.
Look at what their anti-cheating layer actually does and the same shape appears. It watches typing bursts to catch pasting, it watches for style shifts mid-session, and it fingerprints text that reads as model-generated. Every one of those tests where the words came from. None of them tests whether the words are true. Nobody can patch that. A confident wrong answer, typed by a human, at human speed, in a consistent voice, is invisible to all of it.
There's a related tradeoff with the rubric. They publish it, and they're right to, because a scoring system nobody can inspect isn't auditable. Go and read it. It's a more interesting document than it has any right to be, 5 dimensions and 11 criteria with the exact evidence each one wants, and everything I'm criticizing here is visible inside it. But a published rubric is also a study guide, and my own retake is the proof of concept. Transparency and gameability turn out to be the same property seen from 2 sides.
Which brings me back to what I keep telling people in my community. Experience can't be faked, but only in one specific setting. You can have AI build the landing page, generate the graphics, write the case study, produce an entire portfolio. What you can't do is sit across from somebody who knows the work and explain why you made any of those choices. It comes apart there, every time, and it's the one test nobody has automated.
So an AI assessment is a screen, not a verdict. It's a very good first conversation at a scale no human can match, which is worth a great deal and isn't the same as an answer. The final call still belongs to a person watching somebody defend their own work.

Who an AI Fluency Assessment Is Actually For
I'm not going to pretend I use this daily. We're 4 people, and I have written before about why small teams beat large agencies. When I hire, I can still do it by hand, and by hand is better than any tool at that size.
The case for it starts where by hand stops working. Their own numbers put the spread of AI fluency inside a single organization at 82 points, from 15 to 97. That's the gap between 2 people sitting near each other who both wrote "AI proficient" on their internal profile, not a gap between companies. Sort a hundred of those by hand and tell me how it goes.
Since this is the part where I recommend something, let me be plain about where I stand. This isn't a sponsored piece and nobody paid for it. I know the founder from a contract we both worked years ago, which is how the assessment reached me at all, and that's the entire relationship. I took it months before anybody asked me to write anything. I'm writing this because I hit this problem myself trying to scale a team, and because I spend a lot of my week around people building products who are hitting the same wall. That's the alignment, and there's not another one.
So my recommendation is specific rather than enthusiastic. If you're hiring at volume, or working out what AI skill actually exists inside a team you already have, a structured conversational assessment beats a resume line, and beats a self rating by a mile. It's not perfect. The alternatives are a document anybody can generate and a self report that barely correlates with reality at all.
The version I find most persuasive is the one from the other side of the table. I'd rather work at a company that uses something like this, and not because I enjoy being scored. It tells me what the people around me are likely to be like, and that they were picked for something more than a confident sentence on a CV.
Nothing here is 100%. Any assessment is only as good as the people taking it, edge cases exist, and I found a real scoring flaw in my own report and published it anyway. But AI fluency needs to become a real thing we can name and measure, because the alternative is what we have now: everybody claiming it, nobody checking, and an industry arguing about detectors while the actual question goes unasked.
Take the assessment for a spin. One is free, it runs 20 to 40 minutes, and it's worth the time if you're at all serious about how you use AI.
That's a referral link. There's no affiliate payout and no money reaches me: it earns credits I can only spend on more assessments, which means the only thing I get out of you clicking it is knowing you clicked it. I'd rather you took it and had a discussion with me about it than took my word for any of this.
We have more AI than we've ever had. What we're short of is a reliable way to tell how much of it is any good.
With or without my help - I wish you the best.







