Brand Logo

Prospecting field note

How Should an AI Agent Safely Generate Leads? A Quality Inspector's 7-Step Checklist

The question I hear most often from RevOps teams is not 'can an AI agent generate leads?' It is 'how should an AI agent safely generate leads?' The safe answer is boring: use the same acceptance criteria you would apply to a human SDR, then encode them.

My job is quality, not growth. I review outbound campaigns, AI agent configurations, and data integrations before they touch prospects — roughly 200 deliverables a year. I have rejected about a third of first submissions in 2025, usually because the spec was too vague, not because the copy was weak.

Where this checklist fits

Most teams assume the risk is a weird email. The bigger risk is silent decay: stale contacts, inferred addresses, and no audit trail. The sections below are the seven checks I use before an AI agent is allowed to touch prospects. This checklist applies when you are about to hand a list to an AI SDR tool, connect API data enrichment to an outbound workflow, or scale with a parallel dialer and need guardrails. It also applies when a deadline makes verification feel optional.

Step 1: Define a qualified lead before you write a prompt

Most people assume setting up an AI agent starts with a clever first line. In my reviews, it starts with the lead schema. Whether you are using an okki-go AI agent or a custom script, an agent cannot judge whether a record is good unless you define what good means.

Vague instructions such as 'target CTOs at SaaS companies' leave too much room. A usable lead definition includes the title list, industry, employee count, geography, and intent evidence. It also includes exclusions: customers, competitors, or roles that should never receive the sequence.

Quality checkpoint: take three records at random. Can a reviewer explain why each one passed? If the answer depends on someone's memory rather than the record's data fields, the spec is not ready.

Step 2: Verify first; enrich second

An AI agent should not use a guessed email address. If an address was created by matching first name and company domain, it is an inference, not a verified contact. Labeling it as verified because the format looks correct is the kind of shortcut that causes quiet damage.

This is where API data enrichment earns its place. Instead of sending from a static list, the AI agent calls an enrichment service when a lead is being activated. The service returns fields like current job title, company size, technology signals, and email confidence. The agent can then decide whether enough evidence exists to send.

What most people don't realize is that 'verified email' can mean different things. Some verifiers check syntax and domain only. Others check whether the mailbox accepts messages, but they still cannot prove that the person you want is reading it. Treat data sources as probabilities, not truth. Store the source and timestamp for every field.

If your engineering team wants to orchestrate this through code, the okki-go npm package is a natural fit. My approval checklist for any integration asks the same questions: Can I set a timeout per record? Can the agent continue when one enrichment source returns nothing? Can I see which API supplied which field? If the answer is no, the integration is not ready.

Step 3: Separate 'technically allowed' from 'actually allowed'

Consent is not a field you fill in after the agent has already generated a list. It is a gate the agent has to pass, and the human owner has to prove it.

In the UK and EU, GDPR has applied since 25 May 2018, and electronic marketing often includes extra rules such as PECR. In the US, CAN-SPAM has regulated commercial email for years. For voice or SMS, the TCPA and state rules can matter. This is not legal advice, because the rules change and vary by channel. It is the reason your quality process should include counsel before launch.

What I look for under the hood is an audit trail. For each record, can I see the source URL or batch file, the consent status, the suppression list version, and any do-not-contact flag? If a prospect asks one question and the agent then sends a promotional follow-up, that is another quality failure.

Step 4: Gate every message with a human review point

AI is good at writing. It is not good at knowing which claims are true, which customers are referenceable, and which competitor comparisons are permitted. Those decisions belong to humans.

Our content gate allows an agent to generate, but nothing leaves without passing checks for hallucinated facts, unapproved stats, pricing promises, and blocked phrases. The word 'lightweight' sometimes gets past people because it sounds harmless; if no one has measured implementation effort, it is still an unsupported claim.

This is why human-in-the-loop outreach is a feature, not a limitation. A human approves the first emails and the follow-up paths. The AI handles the variation. But the human must check every variant type, not just the first one. In May 2025 I made that mistake; I reviewed one version and assumed the others matched. The automated gate caught a blocked phrase in a second version before it sent. I was lucky.

Step 5: Control channel logic, including parallel dialer settings

An AI agent usually focuses on email, but if your process includes voice, a parallel dialer changes the risk profile. A parallel dialer calls multiple numbers at the same time and connects live answers to available reps. It is not a quality problem by itself; it becomes one when concurrency ignores capacity.

Rule I use: max lines = max reps who can take a live conversation, not max lines in the contract. If one rep is available, nine dropped calls or 'please hold' moments can create a worse impression than no call at all. If the rep is occupied, the agent should not dial another number.

Voice also brings legal boundaries. The TCPA in the US has specific restrictions on autodialing, and suppression lists apply to calls and texts. A safe AI agent does not reason around a 'do not call' request. It suppresses the number and records the request.

Step 6: Build kill switches so the agent can stop itself

Safety in software is not only about what the system can do; it is also about what it cannot do when conditions change. An AI agent needs kill switches.

Many bulk sender guidelines that became prominent in February 2024, including Google and Yahoo's, treat spam complaint rates near 0.1% to 0.3% as a warning range. That does not guarantee delivery. It is a signal to pause and investigate before reputation damage compounds.

At minimum, I expect these rules:

One: if bounce rate or spam complaint rate passes the alert threshold you set with your deliverability team, pause remaining sends.

Two: if one enrichment source fails or times out, mark the record for re-enrichment instead of proceeding with blanks.

Three: if a prospect replies with an opt-out or legal objection, suppress immediately across email, phone, and LinkedIn.

Four: if the launch checklist is not complete by the deadline, do not send a partial list.

That last rule is uncomfortable. When a product launch is approaching, partial data feels better than nothing. I would rather pay for a faster enrichment tier and wait a day. In Q1 2025, we spent an extra $480 to enrich 1,800 records before a deadline because the alternative was sending 300 emails with low-confidence titles. The premium was not for speed; it was for certainty.

Step 7: Audit replies, not just sends

After the campaign starts, your work is not finished. The riskiest time is often after the first replies arrive, because the agent is now in a conversation and may improvise.

If a reply asks about pricing, the agent might generate a specific number no one approved. If a reply says 'we already use vendor X,' the agent might start naming competitors. Those follow-ups need a separate content gate and audit trail.

I also review false positives. A 'nice timing' reply from someone outside your target role is not a win; it is a signal that your lead definition or data enrichment threshold is too loose. Feed that back into Step 1. This is how the checklist becomes a loop.

Common mistakes that still show up

A few patterns come up again and again when I audit AI lead-generation workflows.

One is relying on a single enrichment source. Every data provider has gaps. A waterfall approach or at least a fallback source reduces guesswork.

Two is enriching too early. Data decays quickly. An email or title enriched ninety days ago is not the same as data enriched when the agent activates the lead.

Three is judging a parallel dialer only by call volume. If suppressed numbers and rep capacity are not configured, volume just increases the speed of bad outcomes.

Four is checking only the first email. Follow-ups often contain the highest-risk claims because they are generated in response to a previous reply.

Five is treating no response as permission. Silence is not opt-in. A cap on touches should be part of the campaign spec, not an afterthought.

Bottom line: certainty is a quality metric

At okkigo, I am not trying to build the boldest AI agent. I am trying to build the most predictable one. When someone asks how an AI agent should safely generate leads, my answer starts with quality gates and ends with humans who enforce them.

A safe agent is not slower. It is just harder to fool. Deadlines will always create pressure to skip steps, but every skipped step eventually shows up in a bounce rate, a compliance issue, or a wasted follow-up. The goal is to make the safe path the reliable path, so that 'we need to move fast' and 'this is under control' mean the same thing.

Julian Hartwell

Julian Hartwell

Julian Hartwell is an independent B2B sales intelligence analyst covering contact databases, company data, decision-maker profiles, direct dials, prospect lists, and buying signals. He applies the ISO/IEC 25012 data-quality model while examining field accuracy, coverage, freshness, duplicate rate, match confidence, and source transparency. His evidence-led guides help revenue teams compare prospecting platforms, define acceptable data thresholds, and build account lists that support reliable territory planning and outreach.