The gist
Imagine you could hire someone who, overnight, gathers 500 contacts of potential clients (leads), visits each one's website, reads up on the business, writes a personalized icebreaker (an opening line for a conversation), and has a finished CSV waiting for you in the morning. That's exactly what a Lead Generation System built with Claude Code does. It's full B2B (selling to businesses) automation: from an empty list to an enriched database you can take straight into cold outreach (contacting potential clients you don't know yet).
Key concepts
- Lead Generation Pipeline (a processing assembly line): 4 stages: collect → filter → enrich → export
- Apify: a marketplace of scrapers (programs that collect data), including a Google Maps scraper for finding businesses by niche and location
- Filtering: keep only leads with a website and a phone number (drop the rest)
- Enrichment: Claude analyzes each website: value proposition, what makes them different, a case study (a write-up of a real project with results)
- Output: a CSV with the fields company, location, website, phone, value_prop, differentiator, case_study
- Skill (a reusable Claude module): the whole pipeline gets packaged into a reusable skill with a slash command
Theory
How the pipeline is structured
Input:
niche = "roofing companies"
location = "North Carolina"
count = 30
Step 1: Google Maps Scraping (Apify)
→ 34 businesses with Google Maps metadata
Step 2: Filtering
→ Keep only: website AND phone exist
→ For example: 22 of 34 pass the filter
Step 3: Website Scraping + Claude Enrichment
→ For each of the 22: scrape homepage + /about + /services
→ Claude extracts: value_prop, differentiator, case_study
Step 4: CSV Export
→ leads.csv — ready for Google Sheets and your CRMWhy Apify and not scraping Google directly
Google Maps has an official paid API (application programming interface), but it isn't designed for exporting large business databases. Scraping Google's pages directly violates its ToS (Terms of Service) and gets you blocked quickly. Terms of service change: before you run anything, check the current terms and the laws where you live.
Apify is a marketplace of ready-made scrapers with:
- Proxy rotation (to avoid getting blocked)
- Rate limit management
- Structured output (JSON)
- A free plan with a small monthly credit (current terms: apify.com/pricing)
The Google Maps scraper on Apify takes a query like "roofing companies near North Carolina" and returns a list of businesses with address, website, phone, rating and hours.
The opening prompt (your request to the AI): explaining the task to Claude Code
An important pattern from practice: describe the whole task first; don't jump straight to "create a skill." Let Claude build and test the system first, then package it into a skill:
Hey, today's task is to build a Lead Generation System. Step 1: I give you a niche, a location and a number of results Step 2: you go to Apify and use the Google Maps scraper Step 3: you filter, keeping only leads WITH a website and a phone number Step 4: you visit each website (homepage + about + services) Step 5: Claude extracts: value proposition, what makes them unique, client case studies Step 6: you output a CSV: company_name, location, website, phone, value_prop, differentiator, case_study Ask clarifying questions before you start.
Claude will ask: do you have an Apify API key, which output format do you prefer, which language to use.
Python script: the system architecture
# lead_generation.py
import os
import csv
import json
import anthropic
from apify_client import ApifyClient
APIFY_TOKEN = os.environ["APIFY_API_TOKEN"]
ANTHROPIC_API_KEY = os.environ["ANTHROPIC_API_KEY"]
def scrape_google_maps(niche: str, location: str, count: int) -> list[dict]:
"""Step 1: Collect businesses with the Apify Google Maps scraper"""
client = ApifyClient(APIFY_TOKEN)
run_input = {
"searchStringsArray": [f"{niche} near {location}"],
"maxCrawledPlacesPerSearch": count,
"language": "en",
"countryCode": "us",
}
# Google Maps Scraper on Apify (compass/crawler-google-places)
run = client.actor("nwua9Gu5YrADL7ZDj").call(run_input=run_input)
results = []
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
results.append({
"company_name": item.get("title", ""),
"location": item.get("address", ""),
"website": item.get("website", ""),
"phone": item.get("phone", ""),
"rating": item.get("totalScore", ""),
})
return results
def filter_leads(raw_leads: list[dict]) -> list[dict]:
"""Step 2: Keep only leads with a website and a phone number"""
filtered = [
lead for lead in raw_leads
if lead.get("website") and lead.get("phone")
]
print(f"Filtering: {len(raw_leads)} → {len(filtered)} leads")
return filtered
def scrape_website(url: str) -> str:
"""Step 3a: Collect text from the website (homepage + about + services)"""
import requests
from bs4 import BeautifulSoup
pages_to_scrape = [url, f"{url}/about", f"{url}/services"]
all_text = []
for page_url in pages_to_scrape:
try:
response = requests.get(page_url, timeout=10, headers={
"User-Agent": "Mozilla/5.0 (compatible; LeadBot/1.0)"
})
if response.status_code == 200:
soup = BeautifulSoup(response.text, "html.parser")
# Remove scripts and styles
for tag in soup(["script", "style", "nav", "footer"]):
tag.decompose()
text = soup.get_text(separator=" ", strip=True)
all_text.append(text[:2000]) # Take the first 2000 characters of each page
except Exception:
continue
return " | ".join(all_text)
def enrich_with_claude(company_name: str, website_text: str) -> dict:
"""Step 3b: Claude extracts structured insights from the website text"""
client = anthropic.Anthropic(api_key=ANTHROPIC_API_KEY)
message = client.messages.create(
model="claude-haiku-4-5", # Haiku is enough for structured extraction; current models: the What's current page
max_tokens=500,
messages=[{
"role": "user",
"content": f"""Analyze this text from the website of {company_name}.
Website text:
{website_text[:3000]}
Extract as JSON:
- value_prop: the main value proposition (1-2 sentences)
- differentiator: what makes them unique among competitors (1 sentence)
- case_study: the client's name and the gist of the case if it's on the website, otherwise null
Return only JSON, without markdown.
"""
}]
)
try:
return json.loads("".join(b.text for b in message.content if b.type == "text"))
except json.JSONDecodeError:
return {
"value_prop": "N/A",
"differentiator": "N/A",
"case_study": None
}
def export_to_csv(leads: list[dict], filename: str = "leads.csv") -> None:
"""Step 4: Export to CSV"""
if not leads:
print("No leads to export")
return
fieldnames = ["company_name", "location", "website", "phone",
"value_prop", "differentiator", "case_study"]
with open(filename, "w", newline="", encoding="utf-8") as f:
writer = csv.DictWriter(f, fieldnames=fieldnames)
writer.writeheader()
for lead in leads:
# Make sure every field is present (protects against shifted columns)
row = {field: lead.get(field, "") for field in fieldnames}
writer.writerow(row)
print(f"Exported {len(leads)} leads to {filename}")
def run_pipeline(niche: str, location: str, count: int, output_file: str) -> None:
"""Run the full pipeline"""
print(f"Starting: {niche} in {location}, {count} results")
# Step 1: Scraping
raw_leads = scrape_google_maps(niche, location, count)
# Step 2: Filtering
filtered_leads = filter_leads(raw_leads)
# Step 3: Enrichment
enriched_leads = []
for i, lead in enumerate(filtered_leads):
print(f" Enriching {i+1}/{len(filtered_leads)}: {lead['company_name']}")
website_text = scrape_website(lead["website"])
enrichment = enrich_with_claude(lead["company_name"], website_text)
enriched_leads.append({**lead, **enrichment})
# Step 4: Export
export_to_csv(enriched_leads, output_file)
print("Done!")
if __name__ == "__main__":
run_pipeline(
niche="roofing companies",
location="North Carolina",
count=30,
output_file="leads.csv"
)Real output: what ends up in the CSV
After a run, the system produces a file with data like this:
| company_name | location | website | phone | value_prop | differentiator | case_study |
|---|---|---|---|---|---|---|
| Example Roofing Co | Charlotte, NC | example.com | +1-555-0100 | Full-service roofing with a family approach and financing | Family-owned; years in business listed on the website | Client: full roof renovation finished with no delays |
| ... | ... | ... | ... | ... | ... | ... |
The row above is made up, just as an example. This CSV goes into Google Sheets through File → Import (if everything lands in one column, Data → Split text to columns will fix it).
Packaging it as a skill + slash command
Once the system works, you ask Claude Code to turn it into a skill:
Now that I've gone through this whole pipeline, create a skill called "lead-generation-system" so I can call it as the slash command /lead-generation-system.
Claude creates this structure:
.claude/skills/lead-generation-system/
├── SKILL.md ← description + instructions for Claude
├── scripts/
│ └── lead_generation.py
└── references/
├── requirements.txt
└── sample_output.csvThe skill becomes a slash command by itself: now you can type /lead-generation-system and the system will ask for the niche, location and count. A separate file in .claude/commands/ is no longer required: commands and skills have been merged, and old command files keep working. To make the skill available in every project, put it in your personal ~/.claude/skills/ folder.
Ethical limits
| OK | Not OK |
|---|---|
| Public data from Google Maps | Email addresses of private individuals |
| Public text on company websites | Owners' personal data |
| Business phone numbers from public sources | Aggressively bypassing CAPTCHAs |
| Analyzing public case studies | Data from private sections of websites |
The system collects only what businesses have publicly posted online themselves, which is usually acceptable for B2B outreach. Rules on personal data and email marketing depend on the country (for example, CAN-SPAM in the US, CASL in Canada, GDPR in Europe): check the law where you and your recipients are before sending anything. This is not legal advice.
Practice
Assignment: Lead generation for a specific niche
- Sign up at apify.com and get an API token (the free plan includes a small credit; current terms are on the Apify website)
- Install the dependencies:
pip install apify-client anthropic requests beautifulsoup4 - Ask Claude Code to build a simplified version of the pipeline for a niche you choose:
- "digital marketing agencies in Miami"
- "yoga studios in Austin"
- IT companies in your city (in a niche you understand)
- Run it with count=10 as a test and make sure the CSV is generated correctly
- Open the CSV in Google Sheets and check that all the columns are in the right place
- Ask Claude Code to package the system as a skill
- Bonus: add a scoring column (0-100) where Claude rates the lead's quality based on whether there's a case_study and how detailed the value_prop is
Goal: go from zero to a working lead generation system with real data.
Tools and resources
- Apify: apify.com, a marketplace of scrapers; Google Maps scraper ID:
nwua9Gu5YrADL7ZDj - apify-client:
pip install apify-client, the Python SDK for Apify - beautifulsoup4:
pip install beautifulsoup4, parses website HTML - requests:
pip install requests, HTTP requests to the leads' websites - Google Sheets: Data → Split text to columns, to open the CSV
- Hunter.io: hunter.io, finds emails by domain (there's a free plan with a monthly credit limit; see the website), adds contact details to the pipeline
- Apollo.io: apollo.io, a full B2B contact database with email verification, an alternative to Apify for leads
- LinkedIn Sales Navigator: business.linkedin.com/sales-solutions, targeted search for decision-makers (see the website for trial terms)
- Instantly.ai: instantly.ai, automated email sending with domain warm-up (for when you scale your outreach)
Common mistakes
- Not filtering leads before enrichment. Enriching each lead costs money (an API call to Claude). Without a filter, you spend your budget on companies with no website or phone that you can't contact anyway. Filter BEFORE enrichment.
- Not checking the CSV by hand. Always eyeball the first 10 rows in Google Sheets. Shifted columns, empty fields, duplicates: all of this is easy to catch early, but at scale it ruins the whole database.
- Too large a count at the start. Start with
count=10as a test. If you go straight to 500, you'll spend money on Apify and the Claude API, spend a lot of time on enrichment, and then discover a bug in step 3. Check the cost and data quality on a small run first. - Not connecting leads to outreach. A CSV without action is just a file. This pipeline feeds cold outreach (→ Cold outreach with Claude) and your CRM (→ the Google Sheets table from the same lesson).
How this connects to other lessons
- → Your first clients: start with warm contacts and add lead gen later
- → Cold outreach with Claude: the CSV from this lesson = the input for personalized emails
- → Real monetization cases: the lead gen pipeline as one of the case breakdowns
Key takeaways
A 4-step pipeline: Apify (scrape) → filter (website + phone) → Claude enrichment → CSV. Each step is independent and tested separately.
Build and test the system first, then ask Claude to package it as a skill. A skill made from a working system beats a skill made "out of thin air."
Protect against shifted CSV columns: set
fieldnamesexplicitly and uselead.get(field, ""). This prevents the phone number from landing in the location column.
Next lesson
The mark stays in this browser only and is never sent anywhere. My progress