Audit Rejected Prospects: Is Your ICP Filter Hiding Good Accounts?
TL;DR: Sample rejected accounts, label them independently, separate missing evidence from genuine poor fit, and change one reversible rule at a time. The goal is not to increase pass rate. It is to understand why qualified accounts were rejected without opening the filter to obvious non-ICP records.
Methodology and limitations
This audit is a proposed evaluation workflow, not a published performance benchmark. It uses Unify documentation on Agent questions, response types, guidance, testing, and the company’s description of human-labeled evals. Sample sizes and review cadence must be chosen from your volume, error cost, and available reviewers. No universal false-negative threshold is asserted. Preserve the rejected population and every intermediate label so the audit can be reproduced. Report counts by reason code and segment, but do not interpret a higher pass rate as improvement unless the known exclusions remain protected. The most important output is a defensible policy change with evidence, an owner, and a rollback path. Keep a dated version of the ICP rubric beside the audit so later reviewers can explain whether the system, the data, or the business policy changed.
What is a qualification false negative?
A qualification false negative is an account that the production rule rejects even though an independent review concludes it meets the intended ICP definition. That definition matters. A seller liking the company is not enough. The reviewer needs a written policy for required attributes, acceptable evidence, exclusions, and how to handle missing data. Without that policy, the audit measures reviewer preference rather than model or rule quality.
| Field | What to record | Why it matters | Allowed values |
|---|---|---|---|
| Production decision | Answer, reason code, rule version | Reproduces the exact rejection | Reject or review |
| Evidence available then | Source title, field, date, excerpt | Prevents hindsight enrichment | Present, missing, conflicting, stale |
| Independent label | Reviewer decision before seeing production output | Reduces anchoring | Good fit, poor fit, unknown |
| Error cause | Primary reason for disagreement | Makes fixes specific | Question, evidence, policy, extraction, reviewer |
| Remediation | Smallest proposed rule change | Supports rollback | Rewrite, source, review path, no change |
Start with rejected records, not anecdotes
Export or query the rejected population with the input evidence that existed at decision time. Preserve the production answer, reason code, rule version, and source freshness. Do not enrich the record before the first review because new evidence changes the question. The first label should answer whether the original rejection was justified by the evidence the system had. A second analysis can ask whether better research would have changed the outcome.
Audit sequence
- Freeze the rubric and rule version
- Sample from actual rejected records
- Preserve decision-time evidence
- Collect independent human labels
- Adjudicate disagreements
- Group errors by cause
- Change one reversible rule
- Retest against positive and negative controls
Separate bad fit from missing evidence
The most useful distinction is not simply pass versus fail. Split rejections into confirmed poor fit, missing evidence, conflicting evidence, stale evidence, policy ambiguity, and system or extraction error. Missing evidence should not automatically become a pass. It should enter a research or review path. Likewise, one contrary data point should not outweigh a hard exclusion unless the policy says it should.
| Reason code | Meaning | Correct next step | Unsafe shortcut |
|---|---|---|---|
| Confirmed exclusion | Reliable evidence matches a written exclusion | Keep rejected | Ignore the policy for a strategic logo |
| Missing evidence | A required fact could not be verified | Research or human review | Treat unknown as poor fit |
| Conflicting evidence | Credible sources disagree | Resolve source authority and freshness | Select the convenient source |
| Stale evidence | The fact may no longer describe the account | Refresh the specific field | Re-enrich every field blindly |
| Policy ambiguity | Reviewers apply the ICP differently | Clarify the rubric | Tune prompts before agreeing on policy |
Use independent human labels
Ask reviewers to label without seeing the automated decision when practical. Give them the same rubric and require a short evidence citation. If reviewers disagree, adjudicate the policy before modifying the model. A disagreement may reveal that the ICP is underspecified rather than that the AI is wrong. Keep the original labels, adjudicated label, and rationale so later audits can detect policy drift.
Change the smallest reversible rule
After grouping errors by cause, fix the narrowest cause first. Rewrite an ambiguous question, add a permitted source, clarify a definition, or route an unknown to review. Avoid widening every threshold at once. Re-run the changed rule against both known false negatives and known true negatives. A fix that recovers attractive accounts but also admits obviously excluded ones is not complete.
| Test set | Expected behavior | Failure signal | Response |
|---|---|---|---|
| Confirmed false negatives | Previously missed good-fit accounts are recovered or reviewed | Same error persists | Rewrite the specific question or evidence path |
| Confirmed true negatives | Hard exclusions remain rejected | Excluded accounts pass | Roll back or narrow the change |
| Missing-evidence cases | Unknowns enter a controlled review path | Unknowns become automatic passes | Restore the uncertainty state |
| Conflicting-evidence cases | Decision names source authority and freshness | Decision hides disagreement | Add provenance to the output |
| New holdout records | Performance is consistent outside the repair set | Only repaired examples improve | Investigate overfitting |
Signals that the filter needs review
- A large share of rejections have no evidence
- Reason codes collapse many different causes
- Reviewers disagree on the ICP definition
- The model answer cannot cite the deciding fact
- Recent good-fit wins resemble rejected accounts
- A rule change improves pass rate but weakens hard exclusions
How Unify Agents fit the workflow
Unify Agents can run on company or person records, answer structured questions, use optional guidance, and be tested on examples. The live documentation recommends straightforward questions and places longer explanations in guidance. Unify’s published eval article also explains why overall accuracy can hide important class-level errors and describes human-labeled ground truth. Use those mechanics to make the qualification decision observable, then keep human review for ambiguous or high-cost exclusions.
Turn audit findings into a controlled release
Write a change proposal that names the failing reason code, the evidence behind the diagnosis, the smallest rule change, and the records used to validate it. Keep prompt wording, guidance, sources, and decision logic under version control or in another auditable history. A reviewer should be able to reproduce the old result and the proposed result from the same decision-time evidence. If the change also requires new data, separate the effect of the evidence source from the effect of the rule so the team knows what actually changed.
Release the correction to a bounded audience first. Monitor the distribution of pass, reject, and review outcomes by the same reason codes used in the audit. Inspect individual records when the distribution shifts, because a rate change alone cannot show whether quality improved. Preserve an explicit escape hatch that routes uncertain records to human review and a rollback that restores the prior version. After the release, add newly adjudicated mistakes to the evaluation set, but keep a separate holdout so future changes are not judged only on cases the team has already studied.
Sign up for Unify to put a reviewed outbound workflow into practice.
Frequently asked questions
What is a false negative in AI lead qualification?
It is a rejected account that an independent, evidence-based review concludes meets the written ICP policy.
Should missing data count as poor fit?
No. Missing data is uncertainty. Route it to research or review unless the written policy explicitly defines absence as an exclusion.
How many rejected accounts should I review?
Choose a sample that reflects your rejection volume and high-cost segments. This article does not assert a universal sample size.
Should reviewers see the AI decision?
Hide it during the first label when practical to reduce anchoring, then reveal it for error analysis and adjudication.
How do I avoid overcorrecting the filter?
Change one rule at a time and retest both known false negatives and known true negatives before rollout.
What should happen to ambiguous accounts?
Keep an explicit unknown or review state with the missing or conflicting evidence recorded. Do not force every record into pass or fail.
Glossary
- False negative: a good-fit account incorrectly rejected
- Decision-time evidence: only the information available when the decision ran
- Independent label: a reviewer decision made without anchoring on the production answer
- Reason code: a stable category describing why a record was rejected
- Holdout: records not used to design the correction
- Policy drift: changes in how people interpret the ICP over time

