2009-2014? That's not the old web. Or I'm old. Take your pick.
mattkevan 1 minutes ago [-]
I’m old and it’s definitely not the old web. I consider the old web to be when PNGs were sliced with Fireworks and laid out in tables.
2009-2014 is Web 2.0, back when people still thought social media was a good idea.
Gormo 53 minutes ago [-]
Yeah, I'd say "old web" is early '90s up through about mid-2000s or so. Geocities, Tripod, Angelfire, lots of standalone web forums, the early blogosphere, no social media, etc.
clickety_clack 17 minutes ago [-]
Pre-social media is probably the watermark. That’s what sucked all the content out of the web and into walled-garden platforms.
stillpointlab 1 hours ago [-]
I've just gotten used to this at this point. My first time on the web was somewhere around 1995 I think, although my first time on the Internet was earlier (it was a proxy through some ancient BBS and I'm pretty sure it was using gopher). Even though I was just a kid back then, clearly that makes me old now.
kQq9oHeAz6wLLS 2 minutes ago [-]
Speaking of gopher, I'm low-key hoping that becomes the new place for all the non-bot traffic. Gopher felt magical back in the pre-www days.
benjaminl 32 minutes ago [-]
For me the old web is when people still had homepages. When those went away, the old web died.
drooopy 52 minutes ago [-]
Right? When I think of the old web I think of my Xena and X-Files fan sites on Geocities circa 1997
darknavi 44 minutes ago [-]
I was a few years late to the start (early 90s kid) but I am nostalgic for the vast amount of myfreewebs + dot.tk websites out there.
Hit counters on the front page was mandatory of course.
nkrisc 55 minutes ago [-]
It was still a significantly different era than today. Let’s call it Middle Web. Maybe Late Middle Web.
ajsnigrutin 49 minutes ago [-]
+1 for this
I'd consider the facebook and mega era to be relatively new, the "old web" for me would be the one without centralization around a few giants, the era of random phpbb forums, private websites with "this site is under construction" banners and internet directories to find stuff.
dd8601fn 37 minutes ago [-]
God… I stood up so, so many phpBB instances.
mryall 53 minutes ago [-]
Quite ironic that a link shortener which went offline for a decade or so is now posting about other sites not staying online.
The plot features two American men who stumble upon Brigadoon, a mysterious Scottish village that appears for only one day every 100 years; one man soon falls in love with a young woman from Brigadoon. The show's song "Almost Like Being in Love" subsequently became a standard.
6c696e7578 1 hours ago [-]
It seems putting anything on the web that allows submit is screaming to get spammed these days. Is there any sort of spam filter that's worth using?
z_rho_one 2 hours ago [-]
Remember the good ol' days when we all thought that everything on the web would exist for eternity and over.
2 hours ago [-]
Gormo 52 minutes ago [-]
I mean, it does usually, just not always in its original location.
nephihaha 2 hours ago [-]
No, that's just embarrassing stuff. That stays on the web forever.
SAI_Peregrinus 2 hours ago [-]
Yep, it's Murphy's law of online content. Anything you want to reference later will be gone, with no archive copies. Anything you want deleted will be available forever.
johnnyanmac 57 minutes ago [-]
Interesting blog post comparing the culture and incentives of CEOs in different countries? Completely gone, tried to search it up multiple times to no avail. It's barely 3 years old.
That random 20 year old video of some middle schoolers doing a flying kick and breaking a vending machine? Yeah, just popped up on my feed yesterday.
2 hours ago [-]
mohamedkoubaa 1 hours ago [-]
Especially since the Internet is literally a _messaging_ protocol.
hmhrex 2 hours ago [-]
Purevolume mention made me sad. I miss that community.
tokai 3 hours ago [-]
Am I getting old? 09-14 is not even close to the old web for me. The old web, to me, was back when people still published physical 'phone' books for websites.
acheron 3 hours ago [-]
Seriously. 2009 is several years after everyone was already saying “web 2.0”! That is nowhere near the “old web”.
rdmuser 2 hours ago [-]
The old web doesn't necessarily mean the oldest web. 12-17 years ago was very much an older fairly different era of the web that's worth analyzing even if it's on the younger side of the old web. I can definitively sympathize with your reaction though, it doesn't feel like that era was that long ago yet.
dasil003 2 hours ago [-]
I vaguely recall those, but they were more for normies trying to get online. For me the old web is what I saw when I logged into my university gopher server and saw the advertisement for something called the World Wide Web which I could browse via lynx. Soon enough I got a PPP connection and then Mosaic/Netscape 1.0. However everything after javascript shipped (let alone CSS) is new new new. I'd almost go as far as saying if it doesn't have a tilde in the URL it's not old web... almost...
dd8601fn 25 minutes ago [-]
I’m vaguely remembering tilde username for our public html folders. Was that old Apache behavior?
And now I wonder if Apache is even still common. I spent so much time fiddling with apache configs.
3 hours ago [-]
nipunaeka89 1 hours ago [-]
I was wondering what happened to the old web too
SwellJoe 2 hours ago [-]
"The old web" is 1993 to 2007. It's all been downhill ever since.
shevy-java 3 hours ago [-]
Webpages dying is probably one of the biggest design flaws of the original web.
I am not saying old content needs to be preserved forever, but so much content has factually been lost over time. Old logs from text-based MUDs for instance, even for MUDs that still exist today.
efskap 3 hours ago [-]
It's the great irony of digital media. Copying data is accurate to the bit and is preserved "as-is", but in practice, it requires someone to maintain servers, to care about it. To separate out what is worth preserving.
We hot-linked to all those image hosts because we couldn't imagine them disappearing.
Archive.org had incredible foresight and if it didn't already exist, I'd call such a project a pipe dream.
sumtechguy 2 hours ago [-]
Even Archive.org is rather limited in what it keeps. I know of a very large site that recently disappeared. archive only has part of the web html part of the site. Everything else is either gone or non accessible.
Levitating 2 hours ago [-]
That's not my experience
marginalia_nu 1 hours ago [-]
> It's the great irony of digital media. Copying data is accurate to the bit and is preserved "as-is", but in practice, it requires someone to maintain servers, to care about it. To separate out what is worth preserving.
This is a bit idealized. In practice copying data is not quite accurate (especially in bulk) and bit-rot is a very real phenomenon, both in flight and in storage.
You sometimes encounter it when dealing with files from the early '00s, it's very common to discover a few of them are corrupt, even if they've only ever been copied between harddrives.
tekne 43 minutes ago [-]
Content-addressed storage and error correcting codes mean that one can make bitrot astronomically unlikely with honestly minimal infra investment.
It's copyright that causes anything to disappear from the web IMO -- torrents never die.
EDIT: I am aware that unseeded torrents do in fact die. But it really doesn't take much to seed a whole hard drive's worth of rarely requested data -- this also detects bitrot and so corrects errors automatically if you're not the only copy.
If you are, there's ECC, as well as making another copy.
marginalia_nu 25 minutes ago [-]
There are mitigations in both software and hardware, but most consumer machines, by default, do almost none of that. No ECC RAM, no error correction in the filesystem.
47 minutes ago [-]
cortesoft 3 hours ago [-]
> We hot-linked to all those image hosts because we couldn't imagine them disappearing.
No, we hot-linked all those image hosts because we didn't want to pay to host it ourselves.
EvanAnderson 2 hours ago [-]
...and I had fun replacing images people directly linked from my server with less-- ahem-- savory images.
I enjoyed the emails I got from a couple people who were adamant I "hacked" their site because their "web developer" linked to images on my server... images that now said stuff like "I'm a loser bandwidth thief!", etc. (I never did use really nasty "shock" images-- just taunting stuff.)
There are many TinyMUD logs that were posted on Usenet, still to be found on Google Groups.
However, logging was controversial amongst mudders. It was almost always rude to log a private conversation without knowledge or consent; it was tacky to indiscriminately log while everyone was in the "hangout room" or Rec Room, as it were, and it was also bad form to post logs to Usenet or share them without redacting player names and other things.
But logging was built-in to most clients, and it was possible for server administrators to log (and hypothetically any malware-in-the-middle could log the cleartext, unencrypted TinyMUD TCP streams.) And many nefarious deeds by nasty players were exposed to the light when their logs were posted.
ChadNauseam 3 hours ago [-]
The technology is still in its infancy unfortunately, so there's no way the web could have been based on it, but I think content-addressing is the long-term play. If I click a link, there are some cases where I want the server to respond with a fresh response just for me (e.g. a website showing the weather). But often I just want whatever content was linked to (e.g. a webpage explaining a math content). In the latter case, it would be nice if the link had a hash of the content in it, and 3rd parties could host copies to keep the link working even if the original operator stopped existing.
Gormo 46 minutes ago [-]
That's pretty much how IPFS works.
rcxdude 3 hours ago [-]
It's pretty difficult to avoid without very significant tradeoffs, though. The closest is content-addressable peer-to-peer networks, but these still rely on someone keeping the information around, and they struggle to scale anywhere near as much.
Gormo 47 minutes ago [-]
> Webpages dying is probably one of the biggest design flaws of the original web.
I'd say it's one of the biggest design flaws of the current web, what with more and more content hidden behind paywalls, increasingly restricted WAFs, and rendered client-side via convoluted JavaScript.
Archiving and mirroring of old-style websites, delivered as static HTML, is simple and straightforward. 20 years from now, most web content from ~1996 to ~2015 will still be accessible, but much of today's web content probably won't.
juleiie 13 minutes ago [-]
Old web was kind of dumb anyway. You can put on the rose tinted glasses and feel elite about browsing some shitty site 20 years ago or enjoy the fruits of modern design.
exitnode 3 hours ago [-]
Wow, that is a great domain!
Lord_Zero 2 hours ago [-]
The blog mentions "0.mk's revenue did not cover hosting" but then goes on to implement expensive AI integration. Not counting cost for tokens to do the development.
Also:
> Reply to any 0.mk email and the message lands in a feedback queue the AI reads, triages, and acts on
Is this dangerous? What about jailbreaking AIs and having it delete everyone's account?
Gormo 50 minutes ago [-]
> The blog mentions "0.mk's revenue did not cover hosting" but then goes on to implement expensive AI integration.
The article doesn't seem to describe the cost of the AI solution. It does imply that it is lower than the cost of maintaining and supporting their service manually.
hmartin 3 hours ago [-]
Site got hugged? Is there a torrent?
colenikol2 1 hours ago [-]
Bravo brat
tdx 4 hours ago [-]
I found an old database backup of 0.mk on a disk I had kept.
0.mk started in 2009 as a passion project built by three of us. We worked on it for a few hours each week around our regular jobs. We eventually closed it in 2014 because the revenue (hint: no revenue) could not cover hosting, development, and the constant work of fighting spam and reviewing abuse.
The recovered historical corpus contains 657,607 links. For this analysis, we followed every one of them.
Of the 655,178 links with safe, crawlable targets, 76.7% no longer returned a loading page. After removing repeated destinations, 78.7% of the 492,620 distinct crawlable URLs still did not load. So duplicate links are not creating the result.
I use “did not load” rather than “gone” deliberately. Some URLs returned 403 or 429 and may have blocked the crawler. Pages that returned 2xx or 3xx count as loading even when they now lead to parked domains, login walls, or removed-content notices.
There is one large distortion in the yearly data. A single account created 83,398 URLs pointing to one hostname in 2011. At URL level, 92.5% of that year did not load. Count each hostname once and the result becomes 61.7%, almost identical to 2010 and 2012.
A few things I did not expect:
- 835 restored links point at Facebook’s old photo CDN. None loaded.
- The first link ever shortened was a CSS stylesheet on a WordPress blog.
- Someone shortened localhost on the second day.
- The longest stored URL is 38,753 characters and repeatedly says TRYING_THE_MAXIMUM_URL.
Most users came from one regional online community, so this is not a census of the whole web. It is a record of what that community shared between 2009 and 2014.
I brought 0.mk back to test whether AI can now handle enough development, spam filtering, abuse review, monitoring, and support to make the service sustainable where the original economics failed.
Happy to answer questions about the crawl, the old data, or the rebuild.
hyperionultra 3 hours ago [-]
How did you managed to obtain that domain? Usually single digit or letter domains are “reserved”.
“The name of the .mk domain consists of a minimum of 1 (one) and a maximum of 63 characters.”
temp0826 2 hours ago [-]
ICANN might have set some rule for .com/.net/.org but it's not universal for all tlds
reticulates 2 hours ago [-]
They’re not reserved, pretty easy to obtain if you have money, starting from less than $1k.
levocardia 2 hours ago [-]
[flagged]
CqtGLRGcukpy 2 hours ago [-]
Proof? Using an AI checker isn't accurate.
randomblock1 2 hours ago [-]
The ENTIRE thing is AI generated. I'm not talking about the article. I'm talking about the entire website, the entire "product". https://0.mk/blog/zero-humans
Also it's super easy to tell by looking at it, way too many LLM-isms. No need for a AI checker tool.
Levitating 2 hours ago [-]
> The name was registered in 2009 because it was the shortest URL possible: a zero, a dot, two letters. Seventeen years later the zero means something else.
Did the zero ever mean anything? It's still a 3 character domain regardless if the first character is a zero.
But maybe that's more a measure of my own age and perceptions rather than an accurate representation of the various eras of the internet/web...
2009-2014 is Web 2.0, back when people still thought social media was a good idea.
Hit counters on the front page was mandatory of course.
I'd consider the facebook and mega era to be relatively new, the "old web" for me would be the one without centralization around a few giants, the era of random phpbb forums, private websites with "this site is under construction" banners and internet directories to find stuff.
0.mk, you had one job…
That random 20 year old video of some middle schoolers doing a flying kick and breaking a vending machine? Yeah, just popped up on my feed yesterday.
And now I wonder if Apache is even still common. I spent so much time fiddling with apache configs.
I am not saying old content needs to be preserved forever, but so much content has factually been lost over time. Old logs from text-based MUDs for instance, even for MUDs that still exist today.
We hot-linked to all those image hosts because we couldn't imagine them disappearing.
Archive.org had incredible foresight and if it didn't already exist, I'd call such a project a pipe dream.
This is a bit idealized. In practice copying data is not quite accurate (especially in bulk) and bit-rot is a very real phenomenon, both in flight and in storage.
You sometimes encounter it when dealing with files from the early '00s, it's very common to discover a few of them are corrupt, even if they've only ever been copied between harddrives.
It's copyright that causes anything to disappear from the web IMO -- torrents never die.
EDIT: I am aware that unseeded torrents do in fact die. But it really doesn't take much to seed a whole hard drive's worth of rarely requested data -- this also detects bitrot and so corrects errors automatically if you're not the only copy.
If you are, there's ECC, as well as making another copy.
No, we hot-linked all those image hosts because we didn't want to pay to host it ourselves.
I enjoyed the emails I got from a couple people who were adamant I "hacked" their site because their "web developer" linked to images on my server... images that now said stuff like "I'm a loser bandwidth thief!", etc. (I never did use really nasty "shock" images-- just taunting stuff.)
There are many TinyMUD logs that were posted on Usenet, still to be found on Google Groups.
However, logging was controversial amongst mudders. It was almost always rude to log a private conversation without knowledge or consent; it was tacky to indiscriminately log while everyone was in the "hangout room" or Rec Room, as it were, and it was also bad form to post logs to Usenet or share them without redacting player names and other things.
But logging was built-in to most clients, and it was possible for server administrators to log (and hypothetically any malware-in-the-middle could log the cleartext, unencrypted TinyMUD TCP streams.) And many nefarious deeds by nasty players were exposed to the light when their logs were posted.
I'd say it's one of the biggest design flaws of the current web, what with more and more content hidden behind paywalls, increasingly restricted WAFs, and rendered client-side via convoluted JavaScript.
Archiving and mirroring of old-style websites, delivered as static HTML, is simple and straightforward. 20 years from now, most web content from ~1996 to ~2015 will still be accessible, but much of today's web content probably won't.
Also:
> Reply to any 0.mk email and the message lands in a feedback queue the AI reads, triages, and acts on
Is this dangerous? What about jailbreaking AIs and having it delete everyone's account?
The article doesn't seem to describe the cost of the AI solution. It does imply that it is lower than the cost of maintaining and supporting their service manually.
0.mk started in 2009 as a passion project built by three of us. We worked on it for a few hours each week around our regular jobs. We eventually closed it in 2014 because the revenue (hint: no revenue) could not cover hosting, development, and the constant work of fighting spam and reviewing abuse.
The recovered historical corpus contains 657,607 links. For this analysis, we followed every one of them.
Of the 655,178 links with safe, crawlable targets, 76.7% no longer returned a loading page. After removing repeated destinations, 78.7% of the 492,620 distinct crawlable URLs still did not load. So duplicate links are not creating the result.
I use “did not load” rather than “gone” deliberately. Some URLs returned 403 or 429 and may have blocked the crawler. Pages that returned 2xx or 3xx count as loading even when they now lead to parked domains, login walls, or removed-content notices.
There is one large distortion in the yearly data. A single account created 83,398 URLs pointing to one hostname in 2011. At URL level, 92.5% of that year did not load. Count each hostname once and the result becomes 61.7%, almost identical to 2010 and 2012.
A few things I did not expect:
- 835 restored links point at Facebook’s old photo CDN. None loaded. - The first link ever shortened was a CSS stylesheet on a WordPress blog. - Someone shortened localhost on the second day. - The longest stored URL is 38,753 characters and repeatedly says TRYING_THE_MAXIMUM_URL.
Most users came from one regional online community, so this is not a census of the whole web. It is a record of what that community shared between 2009 and 2014.
I brought 0.mk back to test whether AI can now handle enough development, spam filtering, abuse review, monitoring, and support to make the service sustainable where the original economics failed.
Happy to answer questions about the crawl, the old data, or the rebuild.
“The name of the .mk domain consists of a minimum of 1 (one) and a maximum of 63 characters.”
Also it's super easy to tell by looking at it, way too many LLM-isms. No need for a AI checker tool.
Did the zero ever mean anything? It's still a 3 character domain regardless if the first character is a zero.