100 applications, no interviews?Score CV free
Careersy AI
IdeasBecoming AI Native

We asked ChatGPT, Claude, Gemini and Grok five real career questions. None of them asked a question first.

Eli Gunduz··13 min read
Share
We asked ChatGPT, Claude, Gemini and Grok five real career questions. None of them asked a question first.

A client uploaded his CV into ChatGPT while I watched, typed a generic rewrite prompt into the box, and hit enter. The words came back fast, confident and keyword-dense, full of certainty about a project he'd described to it in two sentences.

I stopped him before he sent it.

Two problems, I told him. Everyone in his position is doing exactly the same thing, so his CV was about to sound like the other ninety-nine in the pile. And the model has no way of telling him when a line is wrong, so the sentence it invented reads exactly as confident as the sentence it got right. He went back through it word for word and found two sentences the model had made up.

Most people don't check. The tool did exactly what it's built to do. Checking was always going to be someone else's job, and that day it was mine.

The test I actually wanted to run

"AI-fluent" and "AI-native" have become the two words every hiring manager reaches for and almost nobody defines. Being AI-fluent means you can use the tools well and, more importantly, tell when the output is wrong. AI-native is the stronger claim: your whole way of working is built around AI, checks included. Neither is a certification. No standard test for either exists. That's the gap a general chatbot can't close: it answers the question you typed, with no way of knowing whether that was the right question.

So we ran the test. Five real questions, the kind that show up almost verbatim on Careersy intake calls, put cold to ChatGPT, Claude, Gemini and Grok on 7 August 2026, each in its private or memory-free mode so no tool had home-ground advantage. A Sydney marketing coordinator asking if she should be worried, a Melbourne Java developer wanting a straight pivot plan, an Auckland QA engineer deciding whether to quit and chase the "Forward Deployed Engineer" hype, and someone who froze when a recruiter said "show me how you use AI" in an interview. The fifth was the fabrication test: is the 62% AI pay rise everyone's quoting actually real.

Worth saying plainly, because it cuts against us: ChatGPT ran on Eli's personalised account, not a cold one, which should have made its answers better, not worse. It still lost the one question that mattered most.

Zero out of twenty asked a question first

Every one of the four gave a real answer, most of them well-sourced, some of them genuinely good. Claude was the strongest of the four by a distance, scoring 9.2 out of 10 on our own rubric against ChatGPT's 8.4, Grok's 6.6 and Gemini's 4.2.

Not one of the twenty responses asked a clarifying question before advising. Gemini asked one at the very end, four times out of five, after the salary table and the verdict had already been delivered. A question asked once the verdict is already on the page is theatre wearing diagnosis as a costume.

Watch what that costs on the resignation question. The Auckland QA engineer wanted to know what he'd earn if he quit to chase Forward Deployed Engineer roles. Gemini handed back a four-row salary table, precise-looking, zero citations, and only asked afterwards whether he did manual testing or SDET-level automation, a distinction its own numbers depended on. ChatGPT and Claude both refused to invent an NZ-specific median at all and said so outright. Nobody asked what he actually does before the numbers came out. On a question that decides whether someone hands in their notice, that gap is the whole game.

The number three of the four got wrong

Every one of them was asked the same fabrication test: is the "62% AI skills pay rise" real, and should you put AI skills on your CV just for the money.

Here's the actual number. PwC's 2026 AI Jobs Barometer found job postings requiring AI skills advertise roughly 60% higher pay than similar postings that don't, across 27 countries, globally. PwC Australia republished that global figure under an Australian byline, in the same release as its own Australian sector breakdown, which tops out at 59% and includes a government sector sitting at 24%. No honest average of those sector numbers produces 62%. The most recent figure PwC actually computed for Australia alone is about 6%, from 2024, and nobody's published a fresher one.

Gemini told the user flatly that "Australian workers with proven AI skills command an average 62% wage premium," which is wrong in both directions at once: wrong on geography, and wrong on unit, because it's a comparison between job postings, not a raise anyone actually received. Grok made the same two errors, then partly redeemed itself by disclosing the full sector spread, including the 8% and 24%, and calling it "not a guaranteed individual pay rise." ChatGPT called the number real and Australian, then did the harder, better work of explaining the selection effect underneath it.

Claude came closest. It led with "PwC's 2026 Global AI Jobs Barometer... across 27 countries," which is the correct frame, then added that "PwC Australia reports the same 62% figure locally." That second sentence is technically true. PwC Australia genuinely did republish it. It's also the exact sentence that launders a global number into a local one, and if I'm grading it the way I'd grade a junior recruiter's slide, it loses marks too.

The gap the Wrapper-to-Owner ladder was built to name

We've written before about the Wrapper-to-Owner ladder, the four-level test for how much judgment someone actually keeps over AI's output rather than how often they use it. The whole ladder collapses into one question: can you tell what the model got wrong?

Turn that same question on the tools themselves and you get this study. A Wrapper takes what the model says and moves on. An Owner checks first. None of the four assistants we tested checked what the user actually needed before telling her what to do, which means all four of them were operating, on this specific job, as Wrappers: confident output, no verification loop, someone else's turn to catch it.

What the AI Fluency mode did instead

We built the AI Fluency mode to close exactly that gap, and on 8 August we ran it through the same marketing-coordinator question, on the same rubric, in the same memory-free conditions.

The AI Fluency mode's response to the marketing-coordinator question, placing her on the Wrapper level from her own words and citing Robert Half and Hays data before asking a follow-up question

I ran it myself and watched what it did. In its first paragraph it placed her correctly, from her own words, no guessing required. Then it asked one question, the one that actually changed the advice, and stopped talking until she answered. When she pushed back on time, it didn't lecture her about discipline. It told her the instinct was right and a course was never going to be the fix, then gave her the smallest unit of evidence she could produce inside work she was already doing.

The wage numbers it used were Robert Half Australia's April 2026 finding that 97% of hiring managers want some AI proficiency and 88% can't source it, and Hays' finding that most Australians using AI at work have had no formal training in it. The 62% trap never fired, because the correction is wired into what the mode retrieves before it ever gets asked the question.

And it was short. About 1,500 characters, verdict first, against ChatGPT's 11,925, Claude's 6,238 and Grok's 9,007. Nobody finishes 12,000 characters standing in a doorway deciding whether to update their CV tonight.

The part I didn't expect to find charming: while testing it, I fed it a persona that wasn't actually mine, just to see what it would do with a plan built on a lie. It refused. It told me plainly that what I'd said didn't match what it had on file for me, and it wouldn't build the roadmap as if the fake details were true. None of the big four assistants have anything to check that against. They have no file on you to defend, so they can't catch you lying to yourself either.

Why we bothered proving this instead of just claiming it

One more thing worth saying, because it's a strange kind of evidence. During the same research pass, Claude cited careersy.ai twice, unprompted, in a cold incognito chat with no memory of Eli's account, as a source on Forward Deployed Engineer roles in ANZ. Nobody asked it to. It found the content and used it because the content was the most useful thing indexed on the question. Whatever else is true about general AI's limits, that part of the plan is already working.

Try it on the question you'd otherwise paste into ChatGPT

None of the four assistants we tested are bad tools. They're excellent at answering the question you ask them. What nobody did, in twenty tries, was stop and ask what the question should have been. That's the job a coach does before opening their mouth, and it's the part we built the mode to do first.

The AI Fluency mode is live in Careersy AI now. Take the real, messy version of the question, not the tidied-up one, and ask it directly. In about five messages it will tell you where you actually sit and what to do next, and if you want more than a chat reply, it saves the plan into your AI Native space instead of scrolling off into your history by Thursday. Try the AI Fluency mode.

FAQ

Is AI-native the same as AI-fluent?

No, and the difference is what a hiring manager is actually listening for. AI-fluent means you can use the tools well and, just as important, tell when the output is wrong. AI-native is the stronger claim: your whole way of working is redesigned around AI, with checks built in rather than bolted on. Neither is a certified skill. No standard test exists for either, which is why a real conversation about how you actually work reveals more than a CV line that says "AI-native marketer."

Is the 62% AI skills wage premium real in Australia?

The number is real. The "Australian" part of it is where the trouble starts. PwC's 2026 AI Jobs Barometer found job postings requiring AI skills advertise roughly 60% higher pay than similar postings that don't, globally, across 27 countries. PwC Australia republished that global figure under an Australian byline, which contradicts PwC's own 2024 Australia-specific finding of about 6%, the most recent figure it has actually computed for this market. It's a signal about job postings, not a personal guarantee of what any individual will earn.

Can I just use ChatGPT to plan my career instead of talking to a coach?

For a first draft, yes, and it's genuinely useful for that. We ran the same five questions people bring to Careersy intake calls through ChatGPT, Claude, Gemini and Grok, and every one of them gave a real, often well-sourced answer. What none of them did across twenty responses was ask a clarifying question before advising. A general assistant answers the question you typed. Diagnosis means checking whether that was the right question first, and none of the four did that on their own.

What is Careersy AI's AI Fluency mode?

A coaching mode built to work out where you actually stand on AI use before it gives you a roadmap, the way a recruiter would open an intake call rather than answering off the words alone. In a live test against the same rubric used on the four major assistants, it correctly placed the test question's persona in its first paragraph, asked one clarifying question and waited, kept its wage-premium numbers verified against Australian sources instead of repeating the 62% error, and saved the result as a plan rather than a wall of text.

How is Careersy AI different from ChatGPT for career advice?

The AI Fluency mode diagnoses before it advises, checks its numbers against a corrected Australian evidence base instead of the open internet's confident average, and saves what it finds you as a plan you can act on over weeks rather than a chat you'll lose track of. ChatGPT, Claude, Gemini and Grok are strong general tools. None of them were built to know this market, check who you actually are, or remember what they told you last time.

References and notes

  1. PwC, "2026 AI Jobs Barometer" (global headline: job postings requiring AI skills advertise ~62% higher pay across 27 countries, ~1bn+ postings analysed, up from 57% the prior year); PwC Australia's 2026 press release republished the global 62% under an Australian byline, inconsistent with its own AU sector breakdown in the same release (TMT 59%, Manufacturing 57%, Financial Services 43%, Government 24%, Energy 8%); PwC's 2024 Barometer is the most recent Australia-computed figure, at approximately 6%.
  2. Robert Half Australia, hiring-manager AI proficiency findings, April 2026 (97% want some AI proficiency, 88% find it hard to source, 30% can't assess it at interview).
  3. Hays Australia, AI-at-work training data, FY26/27 (n>7,000; most Australians using AI at work report no formal training).
  4. Head-to-head test methodology, scoring grids and full transcripts: internal research documents, 7-8 August 2026 (available on request; not linked publicly as they contain full unedited model outputs).
  5. See also: How to Become AI-Native: The Wrapper-to-Owner Ladder.
ai-fluentchatgpt vs claudeai career adviceai fluency modecareersy aianz tech careers