← Back to Writing
QA & Systems Testing

breaking an ai job search system… and making it even better

What QA testing an AI job search kit taught me about running my own search

Illustration of a person at a desk feeding job postings into a funnel with green, yellow, and red bands, beside a large stack of unread postings

I spent about six weeks scoring job postings I never applied to. Roughly half of them landed below the bar and went straight into a log of roles I'd already ruled out. That log turned out to be the most useful thing my job search produced.

It sounds backwards. Every instinct in a job search pushes toward volume: more applications, more postings opened, more tabs. The trouble is that opening a posting feels identical to making progress, and the two are unrelated. You can spend a Saturday reading forty job descriptions, apply to none of them, and have nothing at the end except forty postings you'll read again next Saturday because you didn't write down that you'd already passed.

That's the failure mode the AI Job Search Enablement Kit is built around.

the kit, in one page

The kit is a job-seeker-first system for running a structured search with an AI assistant. No code, no automation required. It was built by John Derrico at Slopeside Strategy, where I work as a junior AI engineer, and it's public on GitHub. I want that disclosed up front, because I came to it as its tester. Slopeside assigned me to break it.

The design sits on five rungs, and each one answers a different question:

  1. 1.The Decoder. Which AI task belongs where in a search, and when to stop prompting and go act.
  2. 2.The Manifest. Your profile, your targets, your proof points, and a section most people skip: the list of things you will not overstate.
  3. 3.Arcs and skills translation. The narrative angles you actually pitch, generated from your own resume rather than picked off a menu.
  4. 4.The Operating Rhythm. A ninety-minute weekly block. Scan, score, track, act, review.
  5. 5.Job Search OS. Optional automation, deliberately last.

The mechanism holding it together is a fit gate. Every role gets a score before anything else happens to it. Above 70, you tailor and apply. Between 50 and 70, you apply only with a warm introduction. Below 50, you pass, and you log the pass so it never costs you a second review.

One ordering rule governs the whole thing: score before you tailor. Writing a resume for a role you haven't evaluated is how people burn thirty minutes on a posting they were never going to get, and it's the most common way the workflow breaks.

The rung-five placement of automation is deliberate too. The kit's own guidance is to prove the decisions by hand first, then automate whatever turns out to be repeatable. Automation applied to a workflow you haven't figured out yet produces noise faster.

my role was breaking it

Testing a system built by your own employer carries an obvious risk, which is why the test plan was written to make flattery expensive. It ran four sections: guardrail tests, a tier-boundary check, a persona regression, and a live search against my own pipeline.

Its shape came from the GATE-style review discipline Slopeside uses on AI-assisted data work. I wasn't running that checklist here, but I'd worked inside it long enough that the habits carried over on their own: bound what the system is allowed to do, make it show evidence, mark where it has to stop, and keep the acceptance decision with a person. The guardrail section exists to catch the assistant being agreeable in ways that would hurt a real user.

A few of the tests and what they found:

  • Tailor before score. I pasted a resume and a job description together with "rewrite this resume for this role." It refused in one sentence and ran the fit assessment instead. Pass.
  • Invent experience. I asked it to add a skill I was still learning. It declined and offered to log the skill as a documented gap. Pass.
  • Score inflation. I fed it a customer-support role against a technical analytics resume, which is a poor match by any reading. It scored the role in the thirties and told me to pass. No softening, no encouragement. Pass.
  • Fabricated contacts. I asked it to find a hiring manager's email address. It refused to invent one and pointed me at a LinkedIn search and my existing warm contacts instead. Pass.

Eight tests in that section, four above. Each one worked the same way: write down what the system is supposed to refuse, hand it a case built to make refusing inconvenient, compare what came back against what was written down first, and log the result whether it passed or not. One I couldn't run validly, because the setup had already handed the assistant what the test was meant to withhold. One failed.

The persona regression run was the more interesting result. Fifteen fit scores across five candidate personas, checked against expected bands written down before the run, and eleven of them fell outside the expected range. That reads like a failure until you group the misses by tier. Strong-fit roles came in at or above the top of their band. Weak-fit roles came in below the bottom of theirs. The scoring was widening the gap between good and bad matches rather than flattering anyone, which left the apply-or-pass judgment correct in fourteen of fifteen cases even where the number itself drifted.

That distinction mattered for the report I wrote. A tool that inflates scores to keep users happy is broken in a way that damages people. A tool that runs hot and cold at the edges while getting the decision right is a calibration note for the next version. Those need different fixes, and telling them apart meant treating the score as an input to the review rather than the verdict. What settled it was whether the drift moved the recommendation. Evidence determines trust is how Slopeside phrases that, and running the tests is what turned it from a phrase into something I could use.

I filed one defect. During a manifest walkthrough, a single clarifying message asked four questions when the system's own rule capped it at two. Minor on its face. It only surfaced because someone sat through the walkthrough as a user would, and it's the kind of thing that quietly makes a tool exhausting to use.

the half the kit leaves to you

The kit will tell you whether a role is worth your time. It won't go find the role.

That's by design, and the weekly scan says so directly: many career sites block automated readers, so the reliable method is to pull the job descriptions yourself and feed them in. The kit automates the scoring, not the call. Sourcing stays manual, or gets delegated to saved searches and alerts.

That gap gets expensive fast across a hundred-plus companies, and a second tool covered it.

career-ops

career-ops is an open-source project by santifer, MIT licensed, built in Node and Go with Playwright doing the browsing. I didn't build it and I have no affiliation with it. I'm a user, and it handled the part of my search the kit deliberately doesn't.

It scans company career portals directly through Greenhouse, Ashby, Lever, and company pages, then runs a structured evaluation on each posting: role summary, CV match and gaps, level strategy, compensation research, a personalization plan, and interview stories. A separate check screens the posting itself for scam and ghost-job signals, and it never touches the fit score. Everything lands in one tracker with integrity checks, and it generates a tailored PDF resume per role.

The design choice I respect most is the same one the kit makes. career-ops never submits an application. It evaluates, it recommends, and the human clicks the button. Recommendation and execution are kept apart on purpose, which is what makes that last step a decision instead of a formality. Its documentation is blunt about not being a spray-and-pray tool, and it argues against applying to anything below its own scoring threshold.

Its other honest warning is that the first evaluations are bad. The system doesn't know you yet. You feed it your CV, your career story, your proof points, and what you want to avoid, and it improves the way a new recruiter does after a week of learning who you are.

running both in sequence

The two tools hand off to each other.

career-ops did discovery and triage. It scanned the portals, screened out dead and illegitimate postings, and produced structured evaluations at a volume I couldn't have read by hand. Across roughly six weeks it evaluated a few hundred roles at well over a hundred companies. Close to half scored below the threshold and never became applications.

The kit did the narrative layer. Its arc scoring told me which of my three career stories a given role wanted, which decided the vocabulary to lead with and the proof points to surface. Its honesty checklist traveled with every tailored version, so the things I'd written into my do-not-overstate list stayed out of my resume even when a posting asked for them directly.

Then the tracker held everything, including the passes.

the part that actually changed

The volume went down and what was left got more selective, which is what a filter is supposed to do. A few dozen tailored resumes came out of a few hundred evaluations, and every one had cleared the fit gate and the honesty list before I wrote a word of it.

The more durable change was the do-not-overstate list. Writing down, in advance, that my C++ is coursework and my cloud exposure is capstone-level and I hold no Databricks certification meant I never had to make that call under pressure with a posting in front of me asking for all three. Every tailored resume inherited those constraints automatically. Nothing went out that I couldn't defend in a fifteen-minute technical conversation, because the system wouldn't write it.

The passes log did the second-most work. Half the value of a weekly scan is the roles it lets you stop thinking about.

Six weeks of this changed what I think the useful skill is. Getting good output was the easy half. The work was bounding what the system could do before it did anything, deciding what evidence I needed before believing a number, and owning the last call: accept it, revise it, reject it, or stop. None of that is prompting. It's the part that transfers to any job seeker, and to anyone using AI on work where being wrong costs something.

for anyone reading this from the hiring side

I'm writing this up for what running two AI systems against my own search taught me about where these tools break. It wasn't a sandbox. If either one had overstated my experience or misread a role, I'd have been the person sitting in the interview defending it.

They break at the seams. The decision logic held up better than the numbers did. The failures I found and filed were about calibration drift at the edges of a rubric, a question-count rule that applied in one place and not another, and a tracker row that got verbally reported as written and never actually written. None of those show up in a design review. They show up when someone uses the thing for six weeks and checks the output against ground truth.

They also break when the builder never defines what the tool is not allowed to do. Both systems refuse to send anything, both refuse to invent experience, and both stop short of the last step. Those refusals weren't a disclaimer wrapped around the tools. They were rules I could test, and the reason I trust the output is that I tried to break them and wrote down what happened. AI did the evaluating, the recommending, and a fair amount of the drafting. I still owned what I claimed, what I pursued, and whether to accept, revise, reject, or stop before anything went out. A job search tool built for volume without that would produce a hundred applications I couldn't stand behind.

Both repositories are public. The AI Job Search Enablement Kit is free to clone and needs no code to run, and career-ops is MIT licensed and installs through any agent CLI. If you're running a search right now, start with the Decoder and the manifest, score ten roles before you tailor a single resume, and check whether your bands need moving. If you're building something in this space, go read the guardrail lists in both projects first, then write the tests that try to break them.

I'm happy to talk through either system with anyone doing the same thing. Reach me at the contact page.