Presslei

How to Find Journalists to Pitch

How I Built a 27,000+ Journalist Database From Scratch

DIY Media Database

Built from scratch without paying for a media list

⌚ 9 min read · 2,000 words

27K+
journalist contacts

Most PR agencies pay $500 to $1,000 a month for media databases like Prowly, Muck Rack, or Cision. We did too, for a few months. Then it clicked: those lists are stale, incomplete, and every other agency is working off the exact same names.

Key TakeawayA list of 500 sharply-targeted journalists with fresh recency data beats a purchased list of 50,000 generic contacts. Quality and recency win over volume, every time.
27,140
Verified contacts in Presslei’s database, each tagged with beat and recency data from real campaigns

“A journalist database isn’t a spreadsheet of email addresses. It’s a living system of relationships, recency signals, and coverage patterns.”

— Salva Jovells, Presslei

Key Takeaway

We built a database of 27,000+ journalists without buying a single media list. Free tools, public data from sitemaps and bylines, LinkedIn exports, and some automation get you there. Any PR team can build a better list than what the paid services sell.

So I built our own, from scratch. It now runs 27,000+ journalists deep, with beat information, contact details, and engagement history attached to each one. Here’s exactly how we did it.

27,000+
Verified Contacts
$0
Media List Cost
5
Data Sources Used
27,140
Verified journalist contacts in Presslei’s database, built from years of active campaign work
60 days
Maximum age of relevant coverage before a journalist’s recency signal degrades
25–50
Optimal contact count for a single campaign, quality over quantity every time
4–8hrs
Time to build a properly researched, verified media list from scratch

Why Commercial Media Databases Fall Short

These tools aren’t useless. They’re convenient. But they have three problems that actually matter:

  1. Everyone has the same contacts. If you and every other agency are pitching off the same Muck Rack list, those journalists are drowning in near-identical pitches. Your “exclusive” story lands next to 50 others using the same database.
  2. The data decays fast. Journalists change beats, switch publications, go freelance. Commercial databases update quarterly at best. By the time you pitch, the contact may have moved on months ago.
  3. They miss the long tail. Freelancers, regional reporters, niche trade press, new hires who haven’t been indexed yet. Some of our best placements came from journalists who weren’t in any commercial database.

Building your own database is more work upfront. But the contacts are fresher, more targeted, and yours alone.

Pro TipUpdate your database monthly. Drop contacts who’ve changed beats or publications, and add journalists who’ve recently covered topics relevant to your upcoming campaigns. A stale database is worse than no database.

Source 1: Mining Your Own Placement History

If you’ve done any PR before, start here. Your past placements are a goldmine of journalist contacts.

We went through over 5,200 historical placements and pulled every byline, email pattern, and publication out of them. That gave us roughly 900 journalists we knew had already covered stories like ours, because they had.

The process:

  1. Export all your placement URLs into a spreadsheet
  2. Visit each article and pull the journalist’s name and any contact info
  3. Note what topics they covered and which pitches they responded to
  4. Cross-reference with LinkedIn to confirm their current role

Yes, this is tedious. We ended up automating most of it with scripts that scrape bylines and author pages. But doing it manually for even your top 100 placements gives you a solid starting list.

The best predictor of a future placement is a past placement. Journalists who covered your kind of story before are significantly more likely to cover it again.

Pro Tip

Personalize every pitch. Reference the journalist’s most recent article and explain why your story matters to their specific audience.

WarningNever send to a scraped email without running it through a verification service first. Unverified addresses mean high bounce rates, and high bounce rates can get your sending domain blacklisted.

DO

  • Start with Google News searches for recent relevant coverage
  • Record the last 3 relevant articles for each journalist
  • Verify email addresses before adding to your outreach list
  • Include freelancers who write for multiple publications
  • Update your database after every campaign with response data

DON’T

  • Buy pre-built journalist databases without verification
  • Add journalists based on publication prestige alone
  • Include journalists who haven’t covered your topic in 90+ days
  • Store journalist data without a legitimate business purpose
  • Skip the LinkedIn employment verification step

This is one of our most effective methods, and I’ve rarely seen another agency talk about it publicly.

The logic: if a journalist wrote about your competitor’s data study and linked to it, there’s a good chance they’ll be interested in your data study on a similar topic.

Here’s the method:

  1. Identify 10 to 15 competitors or similar brands that have earned media coverage
  2. Pull their backlink profiles using Ahrefs, Semrush, or Moz
  3. Filter for editorial links (exclude directories, forums, guest posts)
  4. Extract the journalist bylines from each linking article
  5. Cross-reference against your existing database to find new contacts

When we ran this across 35 PR agency domains, we found 655 scored new contacts that weren’t in any commercial database. These are journalists already proven to cover data-driven PR stories.

We score each contact on the domain rating of the publications they write for, whether we have an email, whether we have a LinkedIn profile, and their geographic region. The highest-scored contacts go straight into our priority outreach queue.

Source 3: Publication Sitemap Scraping

Every news site has a sitemap. Most sitemaps expose author URLs. Those author URLs give up journalist names, beats, and sometimes contact info.

We wrote a script that crawls publication sitemaps, pulls author pages, and extracts byline data. Running it across 60 major publications surfaced over 700 journalist records, roughly 580 of them genuinely new to our database.

The approach:

  1. Find the publication’s sitemap (usually at /sitemap.xml or /sitemap_index.xml)
  2. Look for author-specific sitemaps or URLs containing /author/
  3. Extract names and any associated metadata
  4. Cross-reference with recent articles to identify active beats

This works best on mid-size publications. Massive outlets like the BBC have convoluted sitemap structures, but regional news groups and trade publications are straightforward.

Source 4: LinkedIn Connections

If you’ve been networking in your industry, your LinkedIn connections are an untapped source of journalist contacts.

We exported all 8,100+ LinkedIn connections, filtered for media professionals (editors, journalists, reporters, correspondents), and found 580 media contacts we’d never messaged.

The advantage: these people already accepted a connection request. There’s a baseline relationship. A LinkedIn DM from a connection gets read far more reliably than a cold email.

We keep a dashboard tracking which media connections have been contacted, who responded, and who should be skipped because they explicitly asked not to be pitched. That tracking prevents embarrassing double-pitches.

Key Takeaway

The best pitches answer one question: why should this journalist’s readers care about this right now?

Source 5: Email Pattern Engineering

Here’s where it gets technical. Once you have names and publications, you still need email addresses, and most journalists don’t list theirs publicly.

But email addresses follow patterns. Most publications use one of a handful of formats:

PatternExample
firstname.lastname@john.smith@publication.com
firstname@john@publication.com
firstinitial.lastname@j.smith@publication.com
firstnamelastname@johnsmith@publication.com

We built a pattern map covering 265 publication domains and their specific email formats. Add a new journalist from The Telegraph, Express, or Metro, and we can generate a likely email address instantly.

How we built the pattern map:

  1. Start with journalists whose emails we already know (past correspondence, public bios, etc.)
  2. Extract the pattern for that publication’s domain
  3. Apply it to other journalists at the same publication
  4. Verify using email validation tools

This alone took our email coverage from roughly 25% to over 42% of the database. For a free, DIY approach, that’s a serious return.

The Merge: Deduplication and Scoring

The hardest part of a multi-source database is merging everything without creating duplicates. A journalist can show up in your placement history, your competitor backlinks, and your LinkedIn connections, each time with a slightly different name spelling.

Our deduplication process:

  1. Normalize names (lowercase, strip middle initials, handle hyphenated surnames)
  2. Match on name + publication domain first
  3. Then match on email address for cross-publication matches
  4. Manually review fuzzy matches (similar names at the same publication)

The 27,000+ contacts we have today didn’t come from one clean merge. They’re the result of running these five sources over and over, quarter after quarter, for years, on top of every commercial-database export we’ve cross-checked against along the way. About 42% have verified email addresses. About 26% have LinkedIn profiles. Every contact gets tagged with its source, beat, engagement history, and a priority tier.

Keeping It Fresh

A database is only as good as its last update. A few systems keep ours current:

  • Bounce tracking: every bounced email gets flagged, and we check whether the journalist changed publications.
  • Response tagging: every response (positive, negative, or redirect) gets logged, building a picture of each journalist’s preferences over time.
  • Quarterly re-scraping: we re-run our sitemap and backlink scripts every quarter to catch new journalists and publication changes.
  • LinkedIn monitoring: job-change alerts on key contacts flag when someone switches publications.

Do You Need 27,000+ Contacts?

No. For most campaigns you’re pitching 50 to 80 journalists. A large database just means you can be selective. Instead of pitching everyone who might be relevant, you pitch the 50 people most likely to respond, based on their beat, their publication’s authority, and their past engagement with similar stories.

Start small. Even 200 well-researched, correctly-targeted contacts will outperform a 10,000-name list from a commercial database. The value isn’t in the volume. It’s in the accuracy and the freshness.

Want access to our journalist network for your campaign? When you work with Presslei, you get 27,000+ contacts built from years of research and real campaign outreach. Get in touch.

Frequently Asked Questions

How do you build a journalist database without paying for one?

Start with three free sources: scrape article bylines from publication sitemaps to find active journalists, export your LinkedIn connections to identify media professionals, and mine competitor backlinks to find journalists who cover your industry. Then enrich with email patterns based on each publication’s format.

Are paid media databases worth it?

For most small agencies and in-house teams, no. Paid databases like Cision or Muck Rack charge $5,000 to $15,000 per year and often contain outdated contacts. A self-built database using public data is more current, more targeted, and free.

How do you find journalist email addresses?

Most publications follow predictable email patterns like firstname.lastname@publication.com. Once you identify the pattern for a publication, you can generate emails for any journalist there. Verify them using free tools before sending.

SJ

Salvador Jovells

Founder of Presslei, a reactive digital PR agency based in Zurich. Previously led marketing for two ecommerce brands where he discovered that data-driven reactive PR outperforms traditional approaches by every metric. Connect on LinkedIn.

Salvador Jovells

About the Author

Salvador Jovells

Founder of Presslei. 12+ years in ecommerce SEO across international markets. After a decade of link buying for Hockerty and Sumissura, I reverse-engineered 5,272 earned media placements and founded a reactive PR agency that builds authority through data-driven stories journalists actually want to publish. Based in Zurich.

Related Reading

Ready to earn links instead of buying them?

Get 8–14 editorial placements in top-tier publications. No contracts. No risk. Just results.

Book a Free Strategy Call

$3,000 per campaign · 8–14 links guaranteed · Zero penalty risk

Share this article:
𝕏
in



Founder of Presslei. 12+ years in ecommerce SEO across international markets. After a decade of link buying for Hockerty and Sumissura, I reverse-engineered 5,272 earned media placements and founded a reactive PR agency that builds authority through data-driven stories journalists actually want to publish. Based in Zurich.