Scaling a services business was already the hardest problem in my working life, and hiring A players was already the hardest part of that problem. Then AI landed on top of both, and the hiring pipeline every small company depends on quietly filled with spam.

Some upfront transparency about where we stand: hiring is the biggest challenge at Biz of Dev that we haven't fixed, and we have no recruiters and no hiring team, which is deliberate.

When you're a fresh business trying to hire high value people, outsourcing that judgment is the fastest way to kill yourself right at the start. Before anyone can hire on our behalf, we have to understand our own bar well enough to write it down and run it as a process.

So this piece is that bar, written down, a behind the scenes look at what happens after a resume lands at a small consultancy now that the old process is dead: what we check, what we pay for, what gets you in, and what quietly ends the conversation.

And it's worth saying why this matters more than it looks. Every business runs on 2 engines, how customers arrive and how the work gets delivered, and the vocabulary changes by industry but the engines don't.

Then delivery outgrows your own hands. It stops being a process you designed and becomes the people you picked, which is why hiring sits above every craft skill I have, including the ones I'm actually good at.

The Application Pile Is Mostly AI Now

Post a job opening today and you'll learn the new math fast: roughly 90% of what comes back is spam, candidates spamming AI responses, filling applications with AI, answering your emails with AI, and completing your tests and tasks with AI. Most of them.

The obvious counter is detection, and everyone is busy fighting AI with AI, which doesn't settle anything. Say my detector flags that a candidate used AI. So what?

That alone doesn't mean bad faith, and it doesn't mean the candidate lacks skill either, because AI can be used very poorly and very well, and until I know what they used it for and how, the flag tells me nothing. Whether someone used AI is the wrong question, and how well they used it is the right one, which is the whole argument of my piece on AI fluency.

And in a hiring cycle with a pile of candidates to review, nobody has the time to sit with each one and reconstruct that.

Detection measures whether AI was used and produces witch hunts. Fluency measures how well it was used and produces a skill you can improve.

So right now we lean on 2 channels before a job post. The first is referrals, through a real referral program, because a human vouching with their own reputation is a signal no model can spam.

The second is basically a flywheel, a slower one I'm trying to solidify: Biz of Dev sits inside an ecosystem we run alongside The Wandering Pro and SK NEXUS, and the community around it is full of career minded people, mostly students today, who will graduate and skill up into exactly the talent we want to hire tomorrow.

But that's a preference for this stage, and it won't hold forever. As we grow, our hand will be forced to open roles publicly, and we'll be building a careers space for exactly that.

More of the market hires through referral every month, and I didn't design that, I'm just early about saying it out loud. Either way, a resume still lands on our desk, and the rest of this piece is about what happens after it does, because that part holds whether you were referred or you applied.

The Fundamentals Survived. The Instruments Died.

None of this means the basics stopped mattering. If you're unimpressive on paper, we will still give you an interview shot, but if you're unimpressive on the interview too, sorry, that's the floor everyone has to clear.

A half decent resume that's not full of lies and AI slop (we can tell, faster than you'd think), a real project portfolio, and communication that holds up in a live conversation. Those 3 checks sit in front of everything else, and they filter out most of the pile cheaply.

What died is the instruments we used after the basics. The take home exam is dead, the multi round interview marathon is dead, and every one of them can be faked now, some in an afternoon, some in a prompt.

What can't be faked is actual real paid work, with an OUTPUT at the end of it. That sentence is the whole article, and everything below is just how we operationalized it.

Brief Led Hiring

We call our version Brief Led Hiring, and yes, the name is a deliberate pun on the thing we sell, Discovery Led Product Dev.

Clear the fundamentals with relevant experience and real interest, and we invite you to a project brief: a time boxed problem to solve on an actual product. It's a real problem from a real engagement, anonymized where it needs to be, and paid, because real work gets paid and because payment changes the posture on both sides.

Think of it as a mini probation. Do well, and it extends into standard probation and then a full time seat.

The brief itself is engineered, and the engineering is where 2 years of lessons live:

  • The problem, never the path: the brief states what the product is, what the problem is, what a solved version looks like, and what is out of scope. It deliberately doesn't name the stack, the architecture, or the approach. A named solution is a solution you no longer have to reach, and detail is exactly what a model reproduces best.
  • An assumptions register: we ask what you assumed, why, what would change your mind, and what breaks if you're wrong. There's no faster window into how someone actually thinks.
  • No length limit, on purpose: my instinct was to cap submissions at 2 pages. Bahroze, my cofounder, argued the opposite and won (he was right, annoyingly): cap it, and padded work hides inside the limit looking disciplined. Uncapped, your judgment about what matters is itself part of what we read.
  • Your questions get graded: a requirement question, one only we can answer, is the job being done well. A question Google answers in 10 minutes is a different kind of answer.
  • The sparseness is announced: we say openly that the brief is thin on purpose, so it reads as a design instead of a trick.

Now, one more thing about the money, because it's the part people get wrong. For anything mid level or senior, we don't set the fee, you do: take the brief, scope it, and tell us what it costs.

That single move tells us more than the submission does. How somebody scopes and prices their own work is a signal you can't get any other way, and it arrives before we've spent anything.

Quote too low and you haven't understood the work, quote too high and you haven't read what's out of scope, and if you quote it right with the reasoning attached, we're mostly done deciding.

It also fixes what probation gets wrong. A probation salary is a discount you take because you're uncertain, which means the person you're unsure about is the one paying for your uncertainty.

Unpaid take homes have the mirror problem: nobody is paying, so you can only ask for a few hours, and a few hours of output stopped proving anything the moment a competent looking document became free to generate.

And yes, we do still run unpaid briefs for junior roles, where the volume makes paying everyone impossible, and I'd rather say that plainly than pretend we're purer than we are.

The Walkthrough Is the Test

A document proves nothing now. Anyone can generate a competent looking spec, which means the spec has stopped being evidence of anything, and the document exists for exactly one reason: to give the walkthrough something to interrogate.

So after the brief comes a call, and the call is the filter. Walk us through it, and 3 questions do most of the work.

1 - we pick a load bearing component you specified and ask what it does. 2 - we change the scale of the problem and watch whether you can re-derive your own design. 3 - we look for a step that repeats or contradicts itself and ask why, because generated documents lose track of their own state and a duplicated step is a reliable tell.

The hit rate is the part nobody warns you about. On one recent seat we sent the brief to 7 people and all 7 failed it, and not narrowly either.

They ran the brief through a model, produced a document they couldn't explain, and the walkthrough ended it inside a few minutes.

The 8th came back with the whole thing designed, the integration mapped, an entity diagram, a retry loop, the error cases worked through, all before anybody had paid him a rupee. And he came back with the one thing none of the other 7 produced, which was questions.

That's the tell I trust most now. A model does not come back confused.

And then there's a junior candidate who cleared every fundamental: decent resume, real portfolio, communicated well in the interview. The brief came back, and the work was AI generated end to end, and it didn't make sense.

On the call, they struggled, and struggled really badly. At one point they couldn't answer a question whose answer was written in their own submission, one page below where we were looking.

When your document knows something you don't, the document isn't yours. The entire test, compressed into one exchange.

So understand what we're checking for here, because it's not AI abstinence. We use AI aggressively ourselves, every single day, and the rule is simpler and harder: using AI is fine, and not being able to EXPLAIN what you submitted isn't.

Use AI to research, to draft, to code, to explore approaches you wouldn't have found alone, and then own every line of the result. The candidates who do that are ahead with us, and the ones who can't explain their own submission were never candidates at all, whatever their resume said.

What We Are Actually Measuring

The world has shifted to a place where the output is what matters, and how you get to it matters much less than it used to. If AI wrote your first draft or scaffolded your code, that's fine.

What we care about is that the output clears a bar, and that it gets delivered, internally to the team and externally to the client, through a process that can be trusted. Internally we judge that on 2 buckets, and they overlap rather than splitting 50-50:

  • Adherence: the reliability of you. Are you present in the hours we agreed, on time for clients, updating at end of day, raising blockers out loud instead of going quiet, and doing what you said when you said it. Adherence is measured in the gaps between deliverables, which is why it's the harder one to fake.
  • Quality of Delivery: the caliber of the work. Did you understand the real problem, ask the questions only we could answer, and hand over something correct, reviewed, and explained. Our internal line for the last part: if you can't explain why it works, it's not finished.
How Biz of Dev judges work: Adherence, the reliability of the person, and Quality of Delivery, the caliber of the work, each with 5 observable signals.

And underneath both sits the prerequisite, ownership, which is what produces the other 2.

We run the anti-agency model: a small but mighty team (I've written before about why small teams beat large agencies), taking a client's troubles off their hands and owning everything to the absolute last detail, because we treat clients as partners, never as projects. That posture can't be delegated to a process.

Every person here has to carry it, which is why it's the first thing we look for and the one thing probation can't teach.

Ownership is the prerequisite that produces both Adherence and Quality of Delivery, and the one thing probation can't teach.

The symmetry is brutal and worth stating plainly. You can be genuinely skilled, and if you can't own your work and communicate, you won't make it here. You can own your work and communicate beautifully, and if the delivery doesn't clear the bar, you won't make it here either.

Both halves matter equally, and that's what hiring for output means: output as the proof of both.

So the diagnostic falls out of that. When somebody is failing both halves at once, we stop scoring the 2 axes separately and look underneath, because it's almost never 2 problems.

Skill is trainable. Caring isn't.

There's a phrase in Urdu for the failure mode, phenknay wala kaam karna, and English has no clean single word for it. Throwaway work, doing just enough to push the task off your plate, where the goal was never the outcome, only the removal of the obligation.

A strong engineer who doesn't own the task will hand you a ‘competent looking’ artifact that solves the wrong problem, or a run that stops the exact moment stopping becomes technically defensible.

And it goes unnamed in English, which is probably a good part of why it goes unmanaged.

The Same Test, After You Join

The brief filters for ownership before you're hired. The trouble is that ownership isn't a thing you demonstrate once.

So the board runs backwards here, and nobody gets handed tickets. Every week you post your own plan, the next 3 to 5 tasks, an estimate on each, why each one matters, and where the documentation sits, somewhere the whole team can read it.

The tickets get made from your plan, rather than your plan being assembled out of tickets somebody else wrote.

Anybody who owns a module has that list already, because carrying it is what owning means, and writing it down takes minutes. A ticket taker can't produce it at all, because their model of the job ends at the last ticket they closed, and a generated plan falls apart the moment somebody questions an estimate in front of other people.

What I like about it is that it accuses nobody. It's a cheap habit if you own your work and an impossible one if you don't, so it sorts the team without a performance review anywhere in sight.

Where the A Players Actually Are

For juniors, the number 1 thing we screen for is coachability. Can you be coached, and how much of our time does upskilling you take?

I'd rather hire a coachable junior and build them up than hire a senior who has been spoiled by the market, because that's the sadder story: a lot of senior talent in our market has spent years inside delivery shops and bad agency practice, and the low quality mentality that environment rewards is precisely the thing we're trying to change about this industry.

The 2 traits Biz of Dev screens for when hiring: coachability, the number 1 screen for juniors, and accountability.

And there's an asymmetry nobody likes to say out loud. The people who are senior and actually good mostly aren't looking for jobs, because they have their own business or a side gig running, and capacity finds an outlet.

So the pool that's actively applying skews, structurally, toward people who couldn't do that. It's a rule about pools, and it says nothing about any individual, which is also why, so far, our best people have come from warm channels.

Everyone Here Ends at Partner

Most services businesses around us are delivery shops, project in, project out, and we're deliberately not that.

Project work has real value when done right, but in this era the AI flood is dragging its quality down across the market, and that's not something we're willing to be mixed in with. We're consultants and hands on operators.

That choice shapes the ladder, and the ladder is short by design. People join Biz of Dev as a Product Engineer or a Product Manager, and the end goal, for every single person here, is Product Partner, with a team under them.

Partner is the whole point of the structure, because a consultancy where everyone thinks like a partner is the only version of this business worth building.

The ladder is written down, and we show the whole thing rather than the next rung. That matters.

Somebody 4 months in can't see the slope from inside their first level, so they benchmark us against whatever the market quotes them that week, and the market quotes a starting salary and never a trajectory.

A bigger company will beat our opening number every time. What it can't do is show them the fifth rung, because it doesn't know who they are and won't spend the time finding out.

The cost of showing it is real, and it's worth naming. The moment somebody sees the bands, they negotiate against them, accurately, using my own document, and that has happened.

It's still the right trade, because the alternative is a good person privately wondering whether staying is a mistake, and that wondering is what quietly loses you the people you already trained.

I still do the day to day work myself, alongside the team. Someday that changes, but only when someone as good as me is doing the work that needs doing at that level, and not one day before. I hold myself to that bar too.

The Part I Have Not Solved

Everything above is about getting the right person through the door. The harder problem sits on the other side of it, and I'd rather put our own record on the table than imply we've cracked this.

We're in our second year, and in that time we've hired 4 or 5 people into the team. One is still here (yes, one).

The reason was the same every time: they couldn't think beyond their first level. They could take a ticket and close it, but they couldn't hold a surface, or arrive with a plan, or carry a client relationship.

That isn't a gap another 6 months of coaching closes, because it was never a skill gap.

So the constraint on a business like this is retention. Losing somebody good costs you more than their salary: the engagement context they carried, the client relationship they held, and the year of coaching that got them there, and none of that transfers at handover.

A shop that keeps 5 good people outearns one that cycles 20.

Which means the pitch to keep somebody has to be as deliberate as the pitch to win a client, and built out of things a bigger company structurally can't offer.

A written ladder with a named path to partner. A yearly guarantee I pay out of my own pocket if a project ends. No layoffs, which I can promise because we'll never grow to a size where layoffs become an option.

And one standing rule: if you ever have a reason to leave, tell me first, because there may be something we can work out.

For whatever it's worth as a scoreboard, 2 of the people who left have since messaged asking to come back.

Still Being Figured Out

I want to be straight about the maturity of all this. We're early, the team is small, and this system will evolve as we go.

Working out how to build a consultancy where everyone reaches partner level will probably take me the next 3 to 5 years to get properly right, so what you have read is the current state of an answer, not the final one.

But the core of it won't move: the output is the proof, the brief is the filter, and ownership is the entry fee.

If this read like bureaucracy to you, good news, we just saved each other a lot of time. And if it read like the place you have been trying to find, you now know exactly what happens after your resume lands.

With or without my help - I wish you the best.

Frequently Asked Questions

KNOW WHAT
TO BUILD.

BUILD IT RIGHT.

A quick call to pressure test your idea before you commit resources.

Knowing what problems exist ≠ knowing what to build

Knowing what to build ≠ knowing what to build FIRST

Your insights are valuable. Discovery ensures they turn into the right product, not just a product.

That's exactly how we designed this.

After Discovery, you get:

  • A buildable plan any dev team can execute
  • No proprietary knowledge locked in our heads
  • Exactly what to hire for (if you're hiring)

If you want us to build it, great. If not, we've set you up to succeed with whoever does.

“We're not in the business of trapping clients. We're in the business of making sure you don't waste money - ours or someone else's.”

No. Product Discovery is just as critical for scaling products. Markets change, users evolve, and assumptions expire faster than founders expect.

You get answers across four pillars: Business (what you're solving and how you'll make money), Market (who you're competing against), User (what problems actually matter), and Execution (what to build first and how to validate it). You walk away knowing what to build, what NOT to build, and why - with a plan any team can execute.

CHECK YOUR FIT

Know if Discovery Led Product Dev is right for you in 30 minutes.