Wednesday, March 30, 2011

Just how much should one be unsettled by all this unsettling stuff?

Sitting here as I type this, I can't see you reading it.  With a couple of exceptions, I have little idea who you might be, and you may well know me only through my profile and posts here.  In the absence of face-to-face interaction, it's easy to think oneself anonymous.  For all you know, I could be a dog.

It's easy to think oneself anonymous on the web, but there are significant ways in which this just isn't so. As I've mentioned in a few recent posts, for example, a web server can find out if you're logged in to various sites, your browser quite likely has a unique fingerprint, and the location of your WiFi router is probably in several databases which can be used to locate people who aren't even connected to it (on a recent road trip, I helped install a brand-new wifi router, and did a double-take when Apple's location service knew where I was -- because of the half-dozen other WiFi routers in range).

Which brings us to the question: To what extent are these just somewhat unsettling facts of web.life and to what extent are they cause for real concern?  And of course, the answer is "it depends".

Fine.  What kinds of things does it depend on?

For such things to be more than theoretical concerns, someone has to do something harmful with the information that they wouldn't have been able to do anyway, or at least, the likelihood of someone doing harm has to increase (or if you really want to be technical, the expected harm has to be outweigh any expected benefits, using "expected" in the probability sense).  The downside will depend on what kinds of bad things can happen, which depends on the particular unsettling fact, and how likely they are to happen, which can depend on all sorts of things.

For example, there's probably not a lot of harm in most cases if a server can tell that you're logged into FaceSpace, but if your employer has a strict policy against that and someone decides to install a FaceSpace login detector in your company's internal homepage, the consequences could be serious.  It's up to you to weigh how likely that is and how much you need the job (and how important it really is to browse FaceSpace at work).

If you live under a regime that bans unauthorized WiFi routers, the odds of something bad happening if you put one up anyway are pretty high.  It's almost certainly not worth it.  In most places, however, it shouldn't pose a problem.  If your security is set up properly, most likely all that someone can determine is that there's a WiFi device with some particular SSID near some particular location.  Given that it's easy to determine that there is, say, a house with a mailbox and electricity at a given location, that doesn't seem so dangerous, once you get past the creep-out factor of someone being able to detect something inside your home that they can't see.

To a large extent then, such things are just a part of conducting business on the web.  There's nothing wrong with being concerned about privacy, or taking reasonable steps to try to protect one's privacy, but it's a mistake to expect one's online life to be perfectly private.

But the same can be said of one's offline life.  Ultimately, issues of privacy are not technical, but social and legal.  Expectations of privacy have always been around.  So have breaches of those expectations, and so have various ways of trying to cope with such breaches.  Which is worse: having your Awful Secret shouted in the town square of your medieval village, or to the whole tribal group around the prehistoric campfire, or having it posted to millions of people on the modern web?  I don't see much difference.  All three cases are serious.  What matters isn't how many people find out, but how many people that you care about find out.  I'm not convinced that technology changes that picture much.

Tuesday, March 8, 2011

Wardriving Spartacus

[For those joining recently, "Spartacus" here is a byword for anonymity issues.  See this post, or the anonymity label for more background.]


I was going to respond directly to a comment on the previous post, but thought I'd do a proper post since it ties in to one of the main themes here.

An anonymous  comment (of course) signed "Anli" says:
I like the "Use your router's "MAC cloning" feature". Wouldn't it be nice to have the database of MAC addresses per location? [This would be analogous to a reverse phone directory.  Databases of location per address are widely available.  Inverting one is left as an exercise. -- D.H.]

Oh, well random 48 bits are fine... 
It should be less than 48 if you want to narrow it to a suitable manufacturer...
I started to reply:
Random bits will be fine until the next time the car comes by.  The system will then say, "hmm, don't know this one, let's add it at this location."
All this on a per-database basis.  Skyhook might add you at a different time from Apple, etc.  So yeah, what you want is a database of ...
And then I realized what you really want is a MAC address that's known to be associated with a lot of locations, all over the world, because this is a basic anonymity problem, though with a twist.  It's a basic anonymity problem because the more different locations are associated with your MAC address, that is, the more places you could be, that is, the larger your anonymity set, the less you can be pinned down.

The twist here is that it's very easy to tell if a router at some particular location has the particular MAC address.  By contrast, in the similar-but-different scenario of using an an anonymizer and trying to hide what IP address you're connecting from, we can assume that The Man can tell who's connecting to nodes that are also providing anonymity, but that takes a bit of work -- packet sniffing, etc., and then all they have is a circumstantial case, though perhaps a fairly strong one, that you're participating in or using an anonymizer.

In the case of a wireless router, anyone with $100 worth of parts -- probably quite a bit less, I haven't looked lately -- can tell for sure that there is a router with a given MAC address at a given location.  If The Man in your part of  the world has made it a crime to spoof someone's MAC address, then you can probably expect a knock on the door.

But then, in such a case you probably don't have location services enabled anyway, so why would you be spoofing someone's MAC address?  Likewise, your MAC is unlikely to be in, say, Apple's database, though it will most assuredly be in The Man's.

It's also worth noting that a laptop or phone that's trying to establish its location doesn't need to actually connect to a given wireless router.  It just has to detect packets from it, that is, be within range.  As mentioned previously, the MAC address has to be in the clear for the protocols to work [meaning that you can use WiFi routers to establish your own location without announcing that you're doing it --D.H. Sep 2015].

Summary: If you own a wireless router, expect its location to be known and widely available.  That's not theory.  People do it.

How much of a concern is this, really?  Next post ...

Thursday, March 3, 2011

Now I remember why I don't pay much attention to this kind of stuff

Recently I enabled location services on my MacBook.  That couldn't do too much, right?  The MacBook doesn't have a GPS attached.  A quick check from whatsmyip (go there, then do the "IP Address Lookup") gave a location several miles from my actual one.  Clearly that's as close as it can get.

Well, actually, it was accurate to within a hundred yards or so.

After double-checking that I really, really didn't have a GPS, I dug up what was really happening:   Like many of us, I'm connected to the internet through a wireless router.  That router has a unique MAC address by default.  The MacBook's networking layer knows this (otherwise it can't function), so when I'm on my home network, I'm associated with a particular MAC address.

Several parties, including Apple, keep a database of locations of wireless routers.  They get these locations by "wardriving" -- driving along looking for WiFi signals and using old-school radio techniques to pinpoint roughly where the signal is coming from.  Location services simply contacts the mothership with the MAC address of the router to get the physical location out of the database.  This has all been around a while.  I'm just slow on the uptake.

Assuming everything's working as intended, this doesn't mean that any random person on the internet can find out your location.  You decide whether to share that information, just like you decide whether to share a particular file (caveat: I'm not sure how secure a MacBook's default settings are).  Still, it's another for the growing list of unsettling privacy issues (maybe I'll create a new tag).

As I understand it, there's not a whole lot you can do about this form of information gathering, though I didn't delve deep enough into the IEEE standards docs to be sure I had a definitive answer.  From what I can make out, though, the source and destination MAC addresses have to be on every packet, unencrypted.  Otherwise your router would have to try to decode every encrypted packet it received against the session keys for every active connection in order to see if it was the intended recipient.

Given that, turning off broadcasting of your SSID, using WPA2, whitelisting MAC addresses and so forth is not going to make a difference.  The wardriver just has to sniff for packets and note the MAC addresses.  That's not to say you shouldn't use WPA2 -- you absolutely should, as it provides decent protection against eavesdropping and unauthorized connections -- just that it won't prevent someone from knowing that a router with a given MAC address is in a given location.

There are some countermeasures you can take:
  • Use your router's "MAC cloning" feature to set its MAC to something already in the database (if you choose a MAC not in the database, it will get added with your location next time the car comes by).  Your friends and foes will then think you're in sunny Tahiti or wherever.
  • Don't use WiFi
    • String a bunch of Cat 5 cable and give up the convenience of wirelessness
    • Use a smart phone as a tether -- this has its own set of privacy issues, but I'm not familiar with them and ignorance is bliss [I doubt this helps that much.  A tethering phone looks like any other WiFi router to a wardriver, including having a MAC address.  If you only use the tethering in one place, say your home, there's no real difference from using any other WiFi router.  If you move around, there will be a record that your MAC address was seen at several places, which is probably not what you want --D.H. Sep 2015].
  • Disable location services.  The world will know that there's a wireless router at a given location, but won't be able to associate it with you or your computer (or at least, not quite as easily)
  • Build a suitable Faraday cage around your network.
  • Don't worry, be happy.
Please understand I'm not recommending any of these, except possibly the last.

Tuesday, February 22, 2011

Your browser's fingerprints

This has been out for about a year now, but I just stumbled onto it:

In order to provide a better browsing experience, your browser is prepared to tell any site it visits a number of things about itself.  For example, it may divulge what fonts it knows about, how big your screen is, which browser it is, and so forth.  This is a useful thing to do, and you might not think it would give away much information -- after all, it shouldn't be giving away too much to say you can print Zapf Dingbats and a bunch of other fonts.

Thing is, it gives away quite a bit.  The EFF provides a site, panopticlick, to let you test your own browser setup.  You might not find the results it gives particularly comforting.

In the particular case of fonts, it's not the fonts themselves, so much as the order they appear in.  This turns out to be an artifact of how your particular computer happened to stash the files when you or whoever set up your system installed them, and that's fairly random.  The odds that Zapf Dingbats happens to appear before American Typewriter Condensed Light are closer to 50% than the 0% you might expect assuming the list is sorted alphabetically.  The ordering doesn't seem to be completely arbitrary, but enough so that only a small percentage of browsers out there will actually have the exact same list, taking order into account.

Plugins are even worse.  It's quite possible that only you have your exact combination of plugins, even after sorting (for whatever reason, browsers don't seem to report plugins in a consistent order over time, so the order doesn't provide a stable fingerprint).  I haven't tracked down exactly why this is so, but I believe it's because some plugins are installed on demand as you visit sites.  Which ones you have and haven't collected will depend on which sites you've visited so far, and the exact versions will be affected by when you visited.

Some things that make little or no difference:
  • Whether you have cookies enabled
  • Whether you're using a a normal or anonymous window (e.g., Chrome's incognito feature)
  • In at least some cases, which actual browser you're using -- different browsers may still send the same signature under the covers.
For bonus points, several popular privacy-protection mechanisms can actually make your fingerprint more unique, as they leave traces in the fingerprint and relatively few people use them.  Among them: at least some means of disabling JavaScript (see EFF's paper for details).

Have I mentioned supercookies?

All in all, 80-90% of the browsers that connected to the EFF's site had a unique fingerprint.  Fewer than 1% had a fingerprint shared by more than one other browser (more technically, had an anonymity set with more than two members).  For whatever it's worth, the distribution of fingerprints displays a fine example of a "long tail".

So, yikes.

What to do?  In order of increasing effort:
  • Do nothing.  Roll over, go back to sleep.
  • Do nothing, but bear in mind that browsing is almost certainly not an anonymous activity.  Not a bad assumption in any case.
  • Read the EFF's (and anyone else's) summary of the situation and understand the pitfalls better.
  • Use an anonymizer (but first, read up a bit on the topic -- there are several good links scattered through the posts here tagged "anonymity")
  • Hack your browser to give out less specific information.  Sort those font lists.  Say "version 3.1" instead of "version 3.1.4.1.5.9".
  • Hack your browser to tell randomly varying harmless lies about its setup.  Randomness is important.  Fingerprints will drift naturally over time, but it turns out to be easy to connect a later version (X, Y, Z etc. plus a new plugin) with an earlier one (X, Y and Z).
  • Get the browser developers to change their APIs (e.g., don't give out lists of fonts at all)
  • Get the standards committees to make the underlying protocols more anonymous -- and then get the implementers to implement the standards.
Once you've done all that, sleep soundly knowing that panopticlick was just a proof-of-concept and that people seriously trying to fingerprint browsers have means at their disposal well beyond those mentioned here.

The name panopticlick is a play on Jeremy Bentham's panopticon, a prison design meant to induce "a sentiment of an invisible omniscience" or, in Bentham's own words, "a new mode of obtaining power of mind over mind, in a quantity hitherto without example."

Lovely stuff, that.

Thursday, February 17, 2011

Where Wikipedia pages went to die

While looking for something else (of course) I ran across Deletionpedia, an archive of pages that have been deleted from Wikipedia.  The idea is simple: siphon off pages deleted from Wikipedia, with exceptions such as copyright violations, libel and intentionally offensive pages.

Why do  this?  Wikipedia is reasonably wide-open, but it does have well-known standards for inclusion.  If it's not notable, or contains original research, or creative writing, or anything else that doesn't really belong in an encyclopedia, it's out, regardless of its other merits.  Deletionpedia was an effort to preserve such pages.

I saw "was" because, even though the site is still up, it hasn't been updated since mid-2008 (or 2012, if you believe the rather odd timestamps on the Recent Changes page).  All in all, Deletionpedia collected about 63,000 pages in the space of a few months.  Why did it stop?  The last status update, from 2008, apologizes for recent downtime, promises it will return in improved form and that  "Full service will resume ASAP."

Famous last words, indeed.  Another cool idea that most likely just didn't have sufficient resources behind it, particularly the time required to administer the site and maintain the Python script that was meant to automate the process of sifting out pages that not even Deletionpedia should provide a home for.


The origins of the whole exercise may lie in the "Inclusionist/Deletionist" theological debate in the Wikipedia community.  I wouldn't say that a site like Deltionpedia necessarily supports one side or the other.  On the one hand, it perpetuates pages that would otherwise disappear.  On the other hand, it lowers the consequences of deleting a page.

Neither should such a site have much effect on Wikipedia's "Right to Vanish" which, as far as I can make out, is more of a Right to Make it Somewhat Harder to Associate Your Edits With Your Identity.  Invoking this right does entail deleting one's User: page (but not one's Talk: User page), but I'm not sure how the average user page would make it easier or more difficult to track down who made a particular set of somewhat-anonymized edits.  But I'm not a Wikipedia.expert, so I may have missed something.


Naturally, there is a Wikipedia page on Deltionpedia, and naturally, it has been nominated for deletion at least once.

Sunday, February 6, 2011

You joined the social network ... now see the movie

To be clear right off the bat: This is about the movie The Social Network.  It's not about Facebook, the company, or Mark Zuckerberg, the CEO, or any other actual person, place or thing.  True, there's a person called Mark Zuckerberg in the movie, there is a university called Harvard and a substance called beer, and probably a bit more than the usual amount of care was taken to align those depictions with their real-world counterparts, but it's a movie.  Likewise, my comments here are about the movie.

I liked it.  It's not a bad movie.  But then, I liked Hackers when I finally saw it.

The techspeak is reasonably believable.  In particular, the rapid-fire voiceover as Zuckerberg puts together Facemash is taken directly from the real-life Zuckerberg's online diary (which, however, gets tarted up a bit for the camera).  Using wget to fetch pictures off a web site with an index page full of them is not exactly cutting-edge, but the Zuckerberg character acknowledges as much.  Hacking a perl script with emacs -- or vi, if you prefer -- is a bit more like it.  None of it's neurosurgery, but this is a quick hack.  Judging by the timestamps, he is hacking reasonably quickly, so the guy knows how to get under the hood and get his hands dirty.

What's more interesting is not the coding but the engineering.  In the process of pulling together mugshots of as many Harvard students as he can, Zuckerberg runs across a house whose particular setup makes the task difficult.  What does he do?  Does he down four cans of Jolt Cola and miraculously come up with a superhuman hack to break in?  No.  He punts.

Absolutely the right call.

To make Facemash work, he just needs a bunch of pictures.  He doesn't need every single one, which is fortunate because many aren't online at all.  So why waste time trying to pick the high-hanging fruit when the low-hanging fruit will do?  That little bit of realism actually making it into a big-budget film, even if it goes by so fast you have to think back to realize it happened, and the lack of the usually obligatory thirty-seconds-of-typing-and-the-magic-ACCESS-GRANTED-popup-fills-the-screen scene (yeah, I'm talking about you, Iron Man 2), make a refreshing change, to say the least.

Now, when the site actually goes up and the kids start having fun with it, the resulting traffic apparently brings the Harvard intranet to its knees.  Seriously?  According to the script there were on the order of twenty thousand hits in two hours, if I remember right.  That's about five hits a second, probably more at peak, but not a lot more, and these are fairly small pages -- a couple of mugshot images and some HTML.  It's all going to Zuckerberg's server, and that's not falling over.  The network can't keep up with one Linux box in someone's dorm room?  Sounds like a bit of dramatic license to me.

Similarly, what does the security chief care if some student was snarfing images off the other dorms' servers?  That's not a security threat, it's an annoyance for whoever's administrating those servers and as I understand it a breach of the undergraduate code of conduct.  Judging by the complete mishmash of setups, the sysadmins are probably students themselves, not the university's IT department.  The security guy's job is to keep outside people from causing mischief, and probably to keep everyone from messing with the more sensitive administrative data, particularly grades.  But I digress somewhat.

Actually, one more bit of geekery: Zuckerberg sits, preoccupied, in an OS class while the professor talks about memory management, page tables and such.  Zuckerberg walks out.  The professor taunts him for giving up, at which point Zuckerberg rattles off the answer the professor was looking for.  Except the correct answer was "Sixteen bit virtual address space?  Do what now?  All you've got is 64K and you're going to swap some of it to disk?  It's 2003.  My phone can eat 64K for a light snack."

But starting a site with hundreds of millions of members isn't about coding, nor is it primarily about software engineering in the larger sense.  It's about pulling together the right ideas and getting the word out.  As Zuckerberg points out later, it's also important to have reliable servers, and (as the movie character doesn't mention but the real-life CEO probably would) things like an extensible platform for third parties, but none of that matters if you don't have something of interest running on them in the first place.

Which is why, at least in the movie version, it seems to me that the Winklevoss twins and their partner were more than fairly compensated for their trouble.  Did Zuckerberg deal badly with them by neglecting to mention that he wasn't really working on their site but was in fact working on his own take on a similar idea?  Of course.  Does that mean they invented facebook and he stole it from them?  Not so much.

It's quite clear that the (movie) Winklevosses would have done the site differently.  For starters, "exclusivity" is not a great way to get to a hundred million members.  Nor did they seem to like the look of the site, though that might have been sour grapes.  For all the talk about the first mover advantage -- and one of these days I'd like to have a look at whether such a thing really exists -- MySpace was already around and known to all involved.  For that matter the Harvard house pages that Zuckerberg cribbed called themselves "facebooks".

If Zuckerberg would have been stealing anything, it wasn't the idea of a social networking site, but the Winklevoss's vision of it.   But that wasn't the vision that Zuckerberg implemented.  Facebook (in the movies or real life) isn't just a social network.  It's a collection of features, like relationship status, the wall, the ability to tag photos, a privacy policy, and so forth.  The bulk of these features were implemented well after the split.

The Winklevosses had a concept of a social networking site and they wanted to hire the job done.  They hired the wrong guy and it cost them the few weeks it took them to realize their mistake.  $60 million seems ample compensation for that, even taking into account that their hired hand at least passively misled them.  It's not like Zuckerberg was the only techie in the Cambridge area in 2003 that could have put up a server, or the twins would have had to scrounge for funding to hire someone new.

Again, just going by the movie account.  The real-life version has been hashed out in court.

That's probably deeper into that tar pit than I should have gone.  What's more interesting here is a larger point: Which counts for more: the original idea or the implementation?  There are certainly egregious cases of unscrupulous operators outright stealing an idea and passing it off as their own, but the scenario put forth in the movie isn't such a case.  All other things equal, it's the implementation that counts, just as you can't copyright or patent an idea, only its expression.

At the end of the day, it's not the people with original ideas that tend to go on to business success.  I could rattle off a long list of computing pioneers who didn't become gazillionaires in startups, either because they didn't found startups or the startups didn't succeed.  It's the people who make those ideas into something that people actually use.  Of the mix that goes into that -- design, coding, financing, marketing, knowing the right people (social networking, that is), relentlessness, sheer dumb luck and whatever I left out -- the technical ingredient is arguably one of the most replaceable.

Which brings me back to the mugshot hacking.  The whole hack was nothing but pulling together existing pieces -- the pictures, Apache, wget, perl and yes, emacs to synthesize something new that people wanted.  Nicely foreshadowed.

Thursday, January 27, 2011

Google and Yad Vashem

Organizing the world's information and making it universally accessible and useful means many things.  Among the more solemn:

In observance of Yom HaShoah (Holocaust Remembrance Day), Google and the Israeli Holocaust museum, Yad Vashem, have made available a searchable archive of some 130,000 photos from the museum's collection, as a step toward putting the museum's entire archives online.

[The museum's online digital collections have expanded since this was written.  I changed the link above to the photos collection since the old link was broken.  I'm not sure if it's the same collection or not.  There are several other collections as well, which can be found at https://www.yadvashem.org/collections.html -- D.H. Nov 2018]

OK, this is a bit unsettling ...

File under unintended consequences.  It all makes sense, and yet, it doesn't seem quite right.

Mike Cardwell blogs:
When you visit my website, I can automatically and silently determine if you're logged into Facebook, Twitter, GMail and Digg.
and sure enough, the page will say "Yes, you are logged in" or "No, you are not logged in" at the appropriate places.  Eerie.  What's going on here?

As Cardwell explains, whenever you send an HTTP request to a server, you get back a response code.  That response code might say things like "Your request was OK, here's the data you asked for," or "Sorry, I don't have what you're looking for," or "Goodness, I seem to be having some sort of problem here." or any of a number of other things.  So far, so good.

Modern browsers can keep track of whether you're logged in to particular sites, so you don't have to keep logging in.  Fair enough.  If you're logged in and you ask for something on a site, you'll get it (assuming you have the proper permissions, etc.).  If not, you'll typically get an error.

HTML allows you to reference other web sites within your document -- that's pretty much what makes the web webby -- and modern browsers allow you to behave one way or another depending on what happens when you try to fetch something (it doesn't even have to be based on a status code -- pretty much any reliably observable difference in behavior will do).


Put it all together, and
  • any web site
  • can use a reference to another site
  • to tell if you're logged in to that site
In Chrome, at least, if you open an incognito window to visit Cardwell's site, it can no longer tell whether you're logged in, because incognito windows don't share any state with other browser windows.  But that's kind of throwing out the baby with the bathwater.  You can also turn off JavaScript support (or only selectively turn it on), but that has its own problems.

To really solve the problem you have to be able to control what state is shared between, for example, different tabs or windows.  Doing that simply and non-intrusively is easier said than done.

On the other hand, as a couple of commenters point out, such tricks have been around for a while.  Whether anyone's exploiting them in a significant way is another matter.  Before a site can find out if you're logged in, it has to get you to visit it, not that there aren't plenty of sneaky ways to do that, and then it just knows whether you're logged in or not to sites it knows how to check for (each site requires its own custom-tailored check).  And then, if all you log into is, say, GMail and Twitter, then all your adversary can find out -- from this particular particular, at least -- is that (yawn) you use GMail and Twitter.

Worth losing sleep over?  Probably not.  Worth keeping in mind?  Definitely.

Cardwell's site looks to have a lot of other fun and useful information on it as well ... and if you stop by for a visit, your browser will most likely tell his server I sent you.

The no-tag tag

I recently ran across a blog with a tag I don't think I'd seen before: "No particular tag"

What's the point?  Well, for one thing it gives you an easy way to bring up all the posts that don't have any other tag, and which otherwise couldn't be reached at all through the tag list.

This distinction between nothing and a label for nothing comes up again and again: The empty set vs. no set at all; a null value vs. an empty string or other collection; Odysseus getting Polyphemus to say that "Noman" was attacking him ...

It's a double-edged sword.  It's certainly useful, probably even necessary, to have a something-that-stands-for-nothing, but it can also cause no end of confusion.  Any number of bugs come down to losing track of the distinction between no value and an empty value.

It's a neat idea, adding a tag for no tag, but I'm not sure how much demand there is for it.  If there were much, I'd expect to see more of it.  But perhaps I should leave the definitive statement to the experts:
Everybody knows that more wars have been won with a shovel than a sword. Give a man a hole and what does he have? Nothing, but give a man a shovel and he can dig a hole to contain the nothing.

The not-so-dumb terminal and the cycle of reincarnation

One of the longest-running spectacles in computing is the migration of computing power back and forth between the CPU and its peripherals, particularly the graphics processor:

Start with a CPU and a dumb piece of hardware.  Pretty soon someone notices that the CPU is always telling the dumb piece of hardware the same basic things, over and over.  It would really be more efficient if the hardware could be a bit smarter and just do those basic things itself when the CPU told it to.  So the piece of hardware gets its own computing power, generally some specialized set of chips, to help out with the routine operations.  Just something simple to interpret simple commands and offload some of the busywork.

Over time, the peripheral gets more and more powerful as more functionality is offloaded, and someone realizes that what started out as a few components has effectively become a general-purpose computer, but implemented in an ad-hoc, expensive and unmaintainable fashion.  Might as well use an off-the-shelf CPU.  That works pretty well.  The peripheral is fast, sophisticated and wonderfully customizable.

Then someone notices there are two basically identical CPUs in the system, and people start to write hacks to use the peripheral CPU as a spare, doing things that have little or nothing to do with the original hardware function.  Why not just bring that extra CPU back onto the motherboard and let the hardware device be dumb?

Lather, rinse, repeat ...

With all that in mind, I was going to talk about another prominent cycle, and then I realized that it wasn't really a cycle.  For that matter, the CPU <--> peripheral cycle is only a cycle in the relative amount of horsepower in one place or the other, but even taking that into account ... well, let's just get into it:


Start with a pile of computing power.  It's not much good by itself, so connect something up to it so you can talk to it.  Nothing fancy.  In some of my first computing experiences it was a paper-fed Teletype (TTY) with a 110 baud modem connection to the local computing center.  Later it was a "glass TTY" -- a CRT and a keyboard and a supercharged 2400 baud serial connection to a VAX a couple of rooms over.

Even the dumbest of these CRT terminals could do a couple of things -- clear the screen, display a character, move to the next line -- but not necessarily much of anything more.  But why not?  It's a CRT we're putting characters on, not paper.  We ought to be able to go back and change the characters we've already put up without having to clear the screen and start over.  A couple of improvements, and now you've got a proper video terminal that will let you move the cursor up and down, maybe insert and delete characters, certainly overwrite what's there.

Now, at 2400 baud (about 20 times slower than the "dialup" that everything's faster than), bandwidth is precious, putting pressure on terminal designers to encode more and more elaborate functionality into "escape sequences" -- magic strings of characters that do things like change colors, apply underlines, turn off echoing of characters typed or, if some of the magic characters get dropped, spew gibberish on a perfectly good screen.  For bonus points, let the application actually program macros -- new escape sequences put together out of the existing ones -- getting even more out of just a few characters on the wire.

That's not a glass Teletype any more.  That's a "smart terminal".  Inside the smart terminal is a microprocessor, some RAM for storing things like macro definitions and for tracking what's on the screen, and a ROM full of code telling the microprocessor how to interpret all the special characters and sequences.

In other words, it's a computer.

Well, if it's a computer, it might as well act like one.  Why limit yourself to putting characters on the screen for someone else when roughly the same hardware plus some extra RAM and a disk could do most of the things your wee share of the time-sharing system at the other end of the modem could do?  Thus began the PC revolution that is only just now reaching its endgame.  Sort of [If you think of "PC" as "big boxy thing that sits under your desk", then it's pretty clear PCs are well past their prime.  If you think of "PC" literally as "Personal Computer", we have more of them than ever before -- D.H. Dec 2015].

The problem with cutting the umbilical cord to the central server is that while you may have a pretty useful box, it's no longer connected to anything.  Unless, of course, you buy a modem.  Then you can connect to the local BBS to chat, play games, maybe even transfer some files.

At this point, the box you're talking to may not be particularly more powerful than yours.  Even if you're dialing in to a corporate or university site, there's a good chance that you're still connecting to somebody's workstation, not some central mainframe.  Gone are the days when you connected to "the computer" and it did all the magic.  Now you're connecting your computer to something else it can use.  Relations are much more peer-to-peer, even though there's still a lot of client-server architecture going on.

More importantly, the data has moved outwards.  Instead of one central data store, you've got an ever-growing number of little data stores, which means an ever-growing backlog of routine maintenance -- upgrades, backups, configuration and the like.

If you're using a personal computer at home and you need something that you don't have locally, you have to find the data you need in an ever-increasing collection of places it could be.  If you're a larger institution with a number of workstations, you have the additional problem of making sure everyone sees the same view of important data and configurations.

These basic pressures spur on two major developments: the internet (which is already underway before PCs come along) and the web.

Before too long, things are connected again, except now there are huge numbers of things to connect to, not just one central computer (hmm ... maybe someone could start a business supplying an index to all the stuff out there ...).  With the advent of the web, you have a gazillion web sites all telling a bunch of early-generation browsers the same basic things over and over again.  So ...

... the intelligence starts moving out to the browsers.  Browsers grow scripting languages so that they can  be programmed to respond quickly instead of waiting for instructions from the server.  That cuts out at least three bottlenecks: limited bandwidth, latency between the browser and the server, and the ability of the server to respond to a growing number of connections.  AJAX is born.  Browsers start looking like full-fledged platforms with much the same functionality as the operating system underneath.

On the other hand, data starts moving the other way -- "into the cloud".  For example, email shifts from "download messages to your one and only disk" (POP) to "leave the messages on the server so you can see them from everywhere" (IMAP or webmail).  The more bandwidth you have, the easier this sort of thing is to do, and bandwidth is coming along.  Even so, I'm pretty sure Peter Deutsch will get the last laugh one way or another.


Let's step back a bit and try to figure out what's going on here in broad strokes:
  • From one point of view we have a long cycle
    • In the beginning, all the real work is happening at the other end of a communications link
    • In the middle, all the real work is happening locally
    • These days, more and more real work is happening remotely again -- OK, I haven't run down the numbers on that one, but everybody says it is and I'll take their word for it.
  • On the other hand, today is not a repeat of the old mainframe days
    • A browser is not a dumb terminal.  Even a basic netbook running a minimal configuration has orders of magnitude more CPU, memory and disk than the mainframe of old.
    • There is no center any more.  Even displaying a single web page often involves communicating with several servers in several different locations -- often run by separate entities.
  • The pattern looks different depending on what resource you look at
    • You can make a pretty good argument that data has in fact largely cycled from remote (and centralized) to local to remote again (but decentralized)
    • Computation, on the other hand, has increased all around, and the exact share between local and remote varies depending on the particular application.  I'd hesitate to declare an overall trend.
  • The key drivers are most likely economic
    • Maintaining and administering a bunch of applications locally is more expensive than doing so on a server
    • If bandwidth is expensive relative to computing and storage, you want to do things locally; in the reverse case, you want to do things remotely
Where do we stand with the analogy that started all this, namely the notion that the shift from remote computing to local and back is like the shift of (relative) computing power from CPU to peripheral and back?  Superficially, there's at least a resemblance, but on closer examination, not so much.

Tuesday, January 4, 2011

A strange attractor in web search

Looking through the stats, I see that one of the search terms that landed someone here at Field Notes recently was "How many threes are in a dozen?"

That sentence does appear in this blog, unlikely though that might seem, in a post in which I summarized, among other things, odd search terms that had brought people to Field Notes.  At the time it was sheer coincidence that that search happened to work, but of course since I mentioned it, it's no longer coincidence.  Moreover, since I'm mentioning it again here, I am practically putting myself forth as an expert on the subject.  Search engines are still figuring out the use-mention distinction.

The exact phrase "How many threes are in a dozen?" turns up only two hits (soon to be three).  Since the other one is a discourse on riddles, I should mention that I don't know the intended answer.  The only ones I can come up with are:

  • Um, four, right?
  • None -- there are no threes in "a dozen"
  • 220 (that's the math degree talking)
[And, of course, right after I hit the publish button I realized the right answer is almost certainly 12]