Text & writing
Parts of a URL: Scheme, Host, Path, Query, and Fragment
A URL packs several separate pieces of information into one line. Knowing which part is which explains why some URL differences matter and others do not — useful when comparing links, removing duplicates, or deciding what is safe to delete.
Ready to try the tool this guide describes?
One URL, taken apart
Example: https://shop.example.com:8080/products/shoes?color=red&size=9#reviews
| Part | In the example | What it does |
|---|---|---|
| Scheme | https | How to connect; https is encrypted HTTP |
| Host | shop.example.com | Which server to contact |
| Port | 8080 | Which door on that server; normally left out |
| Path | /products/shoes | Which page or file on the server |
| Query | color=red&size=9 | Extra data for the page, as name=value pairs |
| Fragment | reviews | A position within the page |
Which parts are case-sensitive
- Scheme and host: never. HTTPS://EXAMPLE.COM is the same as https://example.com.
- Path and query: it depends on the server. Many servers treat /About and /about as different pages.
- Fragment: handled by the browser, and case-sensitive when it matches an element on the page.
www is just part of the host
www.example.com is a subdomain of example.com. Most sites send both to the same place, but technically they are different hosts. Links written as www.example.com without https:// are common in print and plain text; browsers add the scheme when you open them.
What the server never sees
The fragment — everything after # — stays in the browser and is not sent to the server. That is why two links that differ only in their fragment load the same page. The query, by contrast, is sent to the server and can change what the page shows.
Percent-encoding
Characters that are not allowed in a URL are written as % followed by two hexadecimal digits: a space becomes %20. Browsers often display the readable form while copying the encoded one, so the same link can look different depending on where you copied it from.
Parts of a URL FAQ
- What is the difference between the path and the query?
- The path names the page; the query, after the ?, passes extra settings or data to it.
- Is the part after # sent to the website?
- No. The fragment stays in the browser and usually scrolls to a section of the page.
- Are URLs case-sensitive?
- The scheme and host are not. The path and query may be, depending on the server.
- Is www.example.com the same as example.com?
- They are different hosts that most sites set up to show the same content.
- Why does my link contain %20?
- It is an encoded space. URLs cannot contain literal spaces, so they are written as %20.
Related guides
- URLs and Trailing PunctuationWhy full stops, brackets, and quotes get attached to links, how good link detection decides what to trim, and how to write links that survive copying.
- UTM and Tracking ParametersWhat utm_source, utm_campaign, gclid, and fbclid mean, which link parameters are safe to remove, and why they make one page appear as many URLs.
Open the tool
Jump into Extract URLs from Text when you are ready to process your files.
