Guides › How to Extract Email Addresses from a Website, Document or Text

How to Extract Email Addresses from a Website, Document or Text

You have a page, a document or a messy export full of text, and you need the email addresses out of it. Doing it by hand is slow and you will miss some. Here are five reliable ways to do it, from a one-click tool to a few lines of code.

The fastest way: paste it into an extractor

Copy the text, paste it into the Email Extractor and the addresses appear instantly, with duplicates removed. Because it runs in your browser, nothing is uploaded. The rest of this guide covers where to get the text from and what to do when addresses are hidden.

Method 1: copy the page source (best for websites)

The text you see on a web page is only part of what is in it. Contact links are usually stored as mailto: links that never appear on screen. To catch them:

  1. Open the page, then press Ctrl+U (Cmd+Option+U on a Mac) to view the source.
  2. Select all (Ctrl+A), copy, and paste it into the extractor.
  3. Repeat for other pages that matter, such as “Contact”, “Team” or “About”.

Method 2: copy from documents and spreadsheets

For Word files, PDFs and spreadsheets, open the file, select everything and copy. For PDFs that are scans, the text is really an image, so run it through an OCR tool first. Paste the text into the extractor. For several files at once, save them as .txt or .csv and drop them onto the input box together.

Method 3: search with a regular expression

Most code editors, including VS Code and Notepad++, can search with regular expressions. Turn on regex mode and search for:

[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Za-z]{2,}

That matches the common shape of an email address: a name, an @, a domain, a dot and an ending. It is a good practical filter, though no short pattern covers every legal address. Use “Find in files” to search a whole folder.

Method 4: a few lines of Python

import re
text = open("page.html", encoding="utf-8").read()
emails = sorted(set(re.findall(r"[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Za-z]{2,}", text.lower())))
print("\n".join(emails))

Converting to lower case and using set removes duplicates. This is handy if you do the job regularly or on many files.

Method 5: handle hidden and obfuscated addresses

Many sites disguise addresses to avoid spam bots. Common tricks, and what to do about them:

Clean the list afterwards

A raw extraction usually needs tidying:

Use the addresses responsibly

Finding an address in public text does not mean you have permission to email it. Laws such as the GDPR (EU and UK), CAN-SPAM (US) and CASL (Canada) restrict unsolicited marketing, and many require consent or a clear opt-out. Use extracted addresses for legitimate purposes, such as contacting people who expect to hear from you or tidying your own lists, and honour unsubscribe requests. This is general information, not legal advice.

Try it now: use the free Email Extractor — private, instant and runs in your browser. Open Email Extractor →

More guides