Tell HN: Amazonbot aggressively scraping my website and ignoring robots.txt

  • Posted 7 hours ago by pera
  • 15 points
At the beginning of the year I decided to set up a scraping and LLM honeypot on one of my personal websites which included a fake git repo with code containing fake HTTP endpoints. The address to this repo was hidden in a public page inside a comment.

About three weeks ago IP addresses from Amazon Searchbot attempted to make requests to the fake endpoints included inside a shell script.

My robots.txt explicitly includes Amazonbot.

I am honestly surprised that this is coming from Amazon. Is this kind of behavior legal?

7 comments

    Loading..
    Loading..
    Loading..
    Loading..
    Loading..
    Loading..
    Loading..