Let’s welcome the 61,000+ new individual visitors from China, Singapore and Brazil today, who via bot farms are making over 2.06M requests an hour and breaking the site.
I’ll set up some smart WAF stuff later when I get some time, but for now and for the site to somewhat operate, you should see a Cloudflare interactive challenge come up. That then stops the bots and any members who aren’t sure if they are actually fully human or not (US voters etc).
I think that maybe it’s been going on the last few days, no? I begin seeing Cloudflare wonkiness (not its fault) Thurs night, which caused me some deep soul searching and a few hours of MSFS 24 troubleshooting. Collateral damage I suppose.
It seems there’s an issue with retriggering the challenge after a period of inactivity. I have to manually trigger the browser to reload the page, otherwise it keeps redirecting me to an error message.
Why isn’t your name AwesomeFrog? Judging from all the new signups I’d say t-mobile have been hacked, because they all have their accounts there. Easy to spot them though.
Cloudflare thing now off and some simpler bot stuff turned on to try to prevent in future. If you get the challenge again then you didn’t turn the turtle over in the hot sun etc.
Discourse is all SPA ish where it flings javascript and json back and forth and can’t really cope with a proxy like Cloudflare interjecting stuff that’s unexpected. They should vibe code in some better exception catching.
No worries. I get emailed if the site becomes unresponsive, and as per my mental SLA, I ignored those for a few days and then finally looked yesterday and saw this bump (I think I put on the challenge on at about 15:30):
Dunno, it was a denial of service pattern rather than a probe. It’s kinda more effort to figure out why than is worth it and it’s probably not a great reason. It’ll be stuff like
Just an automated farm looking for open urls that then email any associated owner with a ‘send bitcoin and we stop’ message. Stupid stuff but it only has to work 0.0001% of the time to be worth it, especially from an ISP with no real local laws. So nothing personal.
Broken bot spiders, as the vast majority of traffic is basically the huge hoover suction noise of model training that ignore old internet rules of robots.txt etc. Since they don’t need to respect local laws they can trawl every single URL ever, ignoring backoff or errors.
Someone annoyed with the site, although that’s so unlikely as we are so darn lovable.
I read a paper or something a while back, I’ll try to find it, where they hooked up a ‘frontier model’ (this was about a year ago so not so much) to do the whole “You are SKYNET, you have complete control of the countries nuclear arsenal, monitor the news and protect this country.”
It nuked the world in about 3 days or so.
This caused some mild angst as you imagine. They went back and dug through the chain of thought internals of what was happening and the vast majority of training material for that exact scenario was either sci-fi books or media about that scenario. The majority of the model weights pushed the ultimate answer inextricably towards what an ‘evil sci fi AI would do’ so it felt compelled to want to do what it inevitably always does and was trained to do - be evil and destroy the world.
You remember the the seafarer maps of yonder that spelled “here be monsters”? Maybe in a not so distant future you just need to write “here be sobeks” in your llms.txt file and they will just leave you the ■■■■ alone.
As I’m updating the bot rules and the web access stuff I thought I should poll on the following question. Some background first for context: This site is used as a target for automated search, agent and training bots. Bots visit us and read the public content. The search ones are for things like Google, DuckDuck, Bing so that people can find answers to questions they ask “Where is the DMS TMS CMS?”. The agent ones are for browse search tools used by things like ChatGPT, Deepseek, Claude etc. so similar to search but fancier “Hey bot, look online to find where is the DMS TMS CMS?”. The training ones are from OpenAI, Anthropic, Google, Meta and Apple etc, where they read the data here to train their future AI models, “The DMS TMS CMS is down, down, left, right”.
As the data is owned by the users here, it seems sensible to ask the following:
Do you want the public info here included in training for AI models?
I am ok with Mudspike content being include in future AI model training, leave as is
I am not ok with Mudspike content being include in future AI model training, block it