Crawler identification and contact page · rockethorseresearch.com
RockethorseResearchBot is an automated web crawler operated by the registrant of this domain. If you found this page from a User-Agent string in your server logs, you are in the right place: this page identifies the crawler, describes how it behaves, and tells you how to reach a person or exclude the crawler from your site.
The crawler gathers publicly available information, primarily product and catalog pages, for private market and product research connected with prospective commercial projects. Retrieved pages are stored and analyzed internally. Page content is not republished, and access to the collected material is not sold or provided to third parties.
robots.txt. It never disguises itself as a browser or another
crawler, and does not rotate identities or proxy addresses.robots.txt, including
Crawl-delay, and including rules addressed to generic groups
(a group such as User-agent: bot is treated as applying to it).
Where a rule is ambiguous, it is interpreted in the site’s favor.403 or
429 response ends the session immediately, suspends crawling
of that site for at least 24 hours, and alerts the crawler’s
operator for review.RockethorseResearchBot (+https://rockethorseresearch.com/bot; contact@rockethorseresearch.com)
To exclude it from your entire site, add this to your robots.txt:
User-agent: RockethorseResearchBot Disallow: /
Path-specific Disallow rules and Crawl-delay are honored
as well. Robots rules are rechecked regularly and a change is normally picked up
within 24 hours; while a site’s rules are cached, the crawler will not
knowingly act against a published change.
contact@rockethorseresearch.com
is a monitored inbox read by a person. Site operators may request exclusion by
email instead of (or in addition to) robots.txt; such requests are
honored and confirmed by reply, normally within ten business days.