http://myfakewebsite.com
http://myfakewebsite.com/articleshttp://myfakewebsite.com/articles/article1http://myfakewebsite.com/articles/article2
http://myfakewebsite.com/jobshttp://myfakewebsite.com/jobs/engineeringhttp://myfakewebsite.com/jobs/design
http://myfakewebsite.com/producthttp://myfakewebsite.com/about
http://myfakewebsite.com with “Follow all links within the domain”.
If you’d like only the articles crawled, better set the URL http://myfakewebsite.com/articles with setting “Only child pages of the provided URL” to ensure the crawler fetches only the pages that contain http://myfakewebsite.com/articles in their URL.
Indexing a single page
If you want only the Engineering page indexed and no other page, you can set the url http://myfakewebsite.com/jobs/engineering combined with the setting “Page Limit” to 1.
Advanced setting: “Depth of Search”
This setting that allows you to say “How many links do I allow the crawler to follow to find a given page?”.
PDF
If your PDFs are stored under the URL you are crawling, they will be included.
Google Docs
If you enter a Google Docs URL, it will be included. Note that if they’re linked from a website, there’s a good chance they are on a different domain, so they won’t be included.
URL must be Public
If a login is required to access the website, Dust will not be able to access its content.
Blocked websites
Some websites are blocking crawling, here is a non exhaustive list:
- reddit.com
- linkedin.com
- instagram.com
- x.com
- tiktok.com