Extract URLs from Text
Pull every web link out of notes, emails, chat logs, or Markdown — one per line, duplicates removed, punctuation trimmed.
Extract URLs from Text works on your text locally in your browser.
Collect every link in one list
Links get buried in research notes, long email threads, chat exports, Markdown files, and documents. Paste the text or upload a document and each web address is pulled out and listed on its own line, in the order it first appears. A full stop or closing bracket that belongs to the sentence rather than the link is trimmed off, while query strings and #anchors are kept intact.
How to extract URLs from text
- Paste your text, or choose Upload Document to open a .txt, .md, .docx, or .pdf file.
- Leave “Include www. addresses” on to also catch links written without http:// or https://, or turn it off.
- The links appear below with counts of links found, unique links, and duplicates removed. Choose Copy Result.
How links are recognized
- A link starts with http:// or https:// — or with www. when that option is on — and ends at a space, a line break, a quote, or an angle bracket.
- A trailing full stop, comma, question mark, exclamation mark, or closing quote is removed; a closing bracket is kept only when the link opened it.
- Everything after ? and # is kept, so parameters and page anchors survive.
- Repeats are removed ignoring capitals in the scheme and domain only; links are listed exactly as written.
| You paste | URLs listed |
|---|---|
| Read https://example.com/guide?page=2#faq. | https://example.com/guide?page=2#faq |
| Sources: (https://en.wikipedia.org/wiki/Mercury_(planet)), www.example.org! | https://en.wikipedia.org/wiki/Mercury_(planet) www.example.org |
| See [the post](https://blog.example.com/p/1) or https://blog.example.com/p/1 again. | https://blog.example.com/p/1 |
Your text stays on your device
- Links are found by JavaScript running in your browser.
- Your text is never uploaded to Looty Tools or any other server.
- None of the links are opened or checked, so no request is made to those sites.
- Uploaded documents are opened and read in your browser; the file itself is never sent to Looty Tools.
Limitations
- Uploads accept .txt and .md files up to 5 MB, Word .docx files up to 15 MB, and PDFs up to 25 MB and 50 pages. Older .doc files are not supported, and scanned or image-only PDFs have no text to use.
- Bare domains such as example.com (no http:// and no www.) are not found, and neither are ftp:, mailto:, or other link types.
- Hyperlinks behind link text in a Word or PDF file are not read — only addresses written out in the text.
- A link that is split across two lines, or that has no space before the next word, is cut at the break or runs into that word.
- Links are not tested, so the list can include pages that no longer exist.
Related text tools
- Remove Duplicate LinesDelete repeated lines and keep one copy of each.
- Sort Lines AlphabeticallySort each line from A to Z.
- Extract Email Addresses from TextPull email addresses out of a block of text.
Guides
- Parts of a URLA URL taken apart piece by piece — scheme, host, port, path, query, and fragment — with which parts are case-sensitive and which never reach the server.
- URLs and Trailing PunctuationWhy full stops, brackets, and quotes get attached to links, how good link detection decides what to trim, and how to write links that survive copying.
- UTM and Tracking ParametersWhat utm_source, utm_campaign, gclid, and fbclid mean, which link parameters are safe to remove, and why they make one page appear as many URLs.
Extract URLs from Text FAQ
- Which links are found?
- Addresses that start with http:// or https://, plus addresses that start with www. (you can switch those off). Bare domains such as example.com, and other kinds of link such as ftp: or mailto:, are not included.
- Why was the full stop after a link removed?
- Punctuation that ends a sentence — a full stop, comma, question mark, closing quote — is almost never part of the address, so it is trimmed. A closing bracket is kept only when the link itself opened one, as in Wikipedia links like …/Mercury_(planet).
- Are query strings and #fragments kept?
- Yes. Everything after ? or # stays, including & and = parameters, so tracking codes and page anchors are listed exactly as they appear.
- How are duplicates decided?
- Two links are the same when they match after ignoring capitals in the https:// part and the domain. Paths and query strings must match exactly, because they can be case-sensitive; http and https versions, or a link with and without a final slash, are listed separately.
- Does it add https:// to www. addresses?
- No. Every link is listed exactly as it was written, so www.example.com stays www.example.com.
- Does it visit or check the links?
- No. Links are found by pattern only; none of them are opened, so the list says nothing about whether a page still exists.
- Can I upload a Word document or PDF instead of pasting?
- Yes. Upload Document accepts .txt, .md, Word .docx, and text-based .pdf files, or you can drop a file onto the text box. The file is read in your browser — it is not sent to Looty Tools — and its text replaces whatever is in the box. Scanned PDFs without a text layer and older .doc files are not supported.
- Is my text uploaded?
- No. Links are found in your browser; your text is never sent to Looty Tools or saved.
