this post was submitted on 24 Jul 2024
182 points (96.0% liked)

Technology

59300 readers
5014 users here now

This is a most excellent place for technology news and articles.


Our Rules


  1. Follow the lemmy.world rules.
  2. Only tech related content.
  3. Be excellent to each another!
  4. Mod approved content bots can post up to 10 articles per day.
  5. Threads asking for personal tech support may be deleted.
  6. Politics threads may be removed.
  7. No memes allowed as posts, OK to post as comments.
  8. Only approved bots from the list below, to ask if your bot can be added please contact us.
  9. Check for duplicates before posting, duplicates may be removed

Approved Bots


founded 1 year ago
MODERATORS
you are viewing a single comment's thread
view the rest of the comments
[–] woelkchen@lemmy.world 3 points 3 months ago (2 children)

test site:reddit.com works fine from DDG for me.

[–] tal@lemmy.today 13 points 3 months ago* (last edited 3 months ago) (2 children)

Older results will still show up, but these search engines are no longer able to “crawl” Reddit, meaning that Google is the only search engine that will turn up results from Reddit going forward.

Robots.txt lets you ask specific user-agents not to index the site. My guess is that that's how they restricted it. I don't know how those changes are reflected in existing indexed pages -- don't know if there's any standard there -- but it'll stop crawlers from downloading new pages.

Try searching for new posts, see how DDG/Bing compares to Google.

EDIT: Yeah. They've got a sitewide ban for all crawlers. That'd normally block Google's bot too, but I bet that they have some offline agreement to have it ignore the thing, operate out-of-spec.

https://www.reddit.com/robots.txt

User-agent: *

Disallow: / 

Here's a snapshot on archive.org's Wayback Machine from April 30 of this year. Very different:

https://web.archive.org/web/20240430000731/http://www.reddit.com/robots.txt

[–] squidspinachfootball@lemm.ee 4 points 3 months ago (1 children)

iirc, isn't robots.txt more of a gentlemen's agreement? I vaguely recall bots being able to crawl a site regardless, it's just that most devs respect robots.txt and don't. Could be wrong though, happy to be corrected.

[–] tal@lemmy.today 4 points 3 months ago* (last edited 3 months ago) (1 children)

Sure, you can write software that violates the spec. But I mean, that'd be true for anything that Reddit can do on their end. Even if they block responses to what they think are bots, software can always try hard to impersonate users and scrape websites. You could go through a VPN, pretend to be a browser being linked to a page.

But major search engines will follow the spec with their crawlers.

EDIT: RFC 9309, Robots Exclusion Protocol

https://datatracker.ietf.org/doc/html/rfc9309

If no matching group exists, crawlers MUST obey the group with a user-agent line with the "*" value, if present.

To evaluate if access to a URI is allowed, a crawler MUST match the paths in "allow" and "disallow" rules against the URI.

EDIT2: Even if, amusingly, Google apparently isn't for this particular case with GoogleBot, given the way that they're signing agreements. They'll honor it for sites that they haven't signed agreements with, though.

EDIT3: Actually, on second thought, GoogleBot may be honoring it too. GoogleBot may not be crawling Reddit anymore. They may have some "direct pipe" that passes comments to Google that bypasses Google's scraper. Less load on both their systems, and lets Google get real-time index updates without having to hammer the hell out of Reddit's backend to see if things have changed. Like, think of how Twitter's search engine is especially useful because it has full-text search through comments and immediately updates the index when someone comments.

[–] squidspinachfootball@lemm.ee 3 points 3 months ago (1 children)

That's a good point, it's probably way less load and overhead if Reddit and Google just sent info back and forth instead of scraping. Good way for Google to keep their spot as the favoured search engine and beat the competition too, since everything that comes up these days are articles full of SEO nonsense at best, then AI generated nonsense at worst. If nobody else can read the actual human responses, Google has a huge leg up. Also interesting to see that Google's honouring the txt file even when nobody's holding them to it.

I had no idea Twitter's search updated their index immediately after a comment is posted though. That's a lot of updates considering the amount of posts they get daily.

[–] tal@lemmy.today 2 points 3 months ago* (last edited 3 months ago)

I had no idea Twitter’s search updated their index immediately after a comment is posted though.

While I never had a Twitter account, it's the major reason that I used the service anonymously. In an unfolding event, like a natural disaster or something, it was absolutely unparalleled in its ability to rapidly comb through enormous amounts of information being plonked in by people around the world. I strongly prefer Reddit-style forum structure most of the time, but for issues for which there is no pre-existing communities and where the common issue is one that will only exist for a short period of time, I think that Twitter's ad-hoc connections between retweets and hashtags works much better than Reddit's association-of-comments-by-subreddit. I understand that Mastodon, unfortunately, doesn't have a full-text search feature, just searching based on exact hashtags. Actually...hmm. I was just talking about Kagi's search lens for the Threadiverse in another comment that I saw. I wonder if Kagi actually indexes Mastodon as well? That'd provide for similar functionality.

investigates

No, it looks like they only do the Reddit-alike Threadiverse (lemmy, kbin, mbin, etc), for which they use the term "Fediverse Forums".

investigates further

It does look like they index in real time, though, or at least quickly -- they probably are one of the institutions out there with an instance slurping up everything out there. I was able to find your comment on that search lens.

That’s a lot of updates considering the amount of posts they get daily.

Yeah, I'm sure that however the Twitter guys built it, they specifically designed it around permitting inexpensive index updates.

[–] eager_eagle@lemmy.world 3 points 3 months ago
User-Agent: bender
Disallow: /my_shiny_metal_ass
[–] itslilith@lemmy.blahaj.zone 1 points 3 months ago* (last edited 3 months ago)

set the date filter to something recent, test site:reddit.com df:w (results from last week only) gives 0 hours hits