Thursday, June 28, 2012

Yet another wacky security scheme

Passwords are easy to get wrong.  Trying to make people come up with "stronger" passwords just makes it worse.  Security questions just provide another avenue of attack, probably an easier one.  So, ladies and gentlemen, may I introduce to you: The security word.

"What is it?", you may later regret asking.

You give the site a "security word".  Later, they will ask you not for the word, but a few randomly selected letters, for example the second, fifth and eighth, and next time it might be the first, fifth and sixth (note to self -- lopado­temacho­selacho­galeo­kranio­leipsano­drim­hypo­trimmato­silphio­parao­melito­katakechy­meno­kichl­epi­kossypho­phatto­perister­alektryon­opte­kephallio­kigklo­peleio­lagoio­siraio­baphe­tragano­pterygon may not be the best choice for this exercise).

If you picked, say, security, and the system asks for the second fifth and eighth letters, you would give 'e', 'r' and 'y'.  If someone's looking over your shoulder, how much information do they have?  Let's fire up the old UNIX shell

$ grep '^.e..r..y.*' /usr/share/dict/words | wc -l
   84

What this means is that there are 84 words in the dictionary on my system that have 'e', 'r' and 'y' in those positions, or about six bits of entropy.  Most of them are words like ventrohysteropexy and dextrogyratory that people are unlikely to pick.  The person who helped me set up the account in question recommended something "easy to remember".  Odds are it's "security".

If not, all an attacker has to do is guess the letters that the site asks for next time.  There's a good chance that at least one will be one the attacker has already seen.  There won't be a lot of choices for the unknown letters.  Without looking at the list, I'd bet that 'q' isn't on it and 'e', 't' and a few others cover most of the possibilities.  Even without having looked over your shoulder, an attacker would know just from the security word being English that certain letters are better to try in certain positions.

So basically you have another hoop to jump through that adds minimal actual security, but tries to create the illusion of strong security, while really just making the system harder to use.  Huzzah.

Wednesday, June 27, 2012

Where's the fire?

It's been a busy fire season in the Southwest, with hot, dry weather and an abundance of fuel.  While I'm generally skeptical about the world-changing potential of social media (ultimately, of the idea that a new technology is necessarily likely to have a great impact), that doesn't mean social media can't play a role.

One example is this fire map from esri (makers of geographic information systems (GIS) software).  Along with official data such as wind information from NOAA and fire perimeters from USGS, it includes layers for Flickr, Twitter, and YouTube activity.  The YouTube feature for some reason got stuck on the same video until I reloaded the page, at which point it worked fine.

The Twitter layer seems to be tagged by the location of the person tweeting.  For example, someone in Fort Collins tweeting a Denver Post story about the Waldo Canyon fire over a hundred miles away is shown in Fort Collins, not where the fire is.  To be fair, it's a lot easier to locate the person tweeting than to figure out from the contents of the tweet that it's really about something somewhere else.

All in all, esri's map seems useful to me mostly for the official information.  The social layer is interesting, but by no means essential.  Searching for "Colorado fire" on Twitter search turns up many more tweets, at least as relevant as those from the map.  Likewise for a YouTube search.  Neither of these searches directly maps the location of the footage, but this doesn't seem like a great obstacle.  Wildfires are quickly given distinctive names ("High Park fire", "Waldo Canyon fire") and you can easily search on those.

And of course wildfires would quickly be given distinctive names.  People need to tell them apart.  If I live in Colorado Springs, I don't care much about the High Park fire, but I care a lot about the Waldo Canyon fire.  As a side effect, it's easy to search for information about a given fire without consulting a map.

And what does such a search find?  Among other things, quite a few links to, and videos from, local newspapers and TV stations.

In short, what does a social-media-enhanced map and search space look like?  A fair bit like one without social media, at least in the context of events with wide interest where there are well-developed traditional media sources.

But broadcasting was never really supposed to be the strong point of social media anyway.

Tuesday, May 1, 2012

Since when are we all on a first-name basis?

Time was, if you had to line a bunch of people up alphabetically, you did it by surname:
  • Ball, Lucille
  • Marx, Harpo
  • Zoolander, Derek
More and more, though, we seem to be arranging people by first name
  • Derek Zoolander
  • Harpo Marx
  • Lucille Ball
Seems odd.  I blame the web.  Gmail does it.  Facebook does it.   Wikipedia can generally autocomplete from a first name ("Ab" gets me Abraham Lincoln at the top) but not the last (spelling out "Lincoln" in full gets me Lincolnshire and a longish list of other names, but nothing on the president).  There are plenty of other examples, I claim.  Indeed, it almost seems to be becoming the norm.

Objectively, there's probably not much to pick between the two schemes.  Without actually measuring, I'd guess that last names tend to be more unique (or "more nearly unique", if you must), but on the other hand we tend to think of people with their first names, well ... first.

Now perhaps this is just a generational thing.  Kids These Days, after all, have no regard for propriety and convention, just like my generation before them.  Perhaps the major software companies have had a hand, spreading their Silicon Valley disregard for regimented old-school thinking.

However, I think Wikipedia's autocomplete is notable.  If we're searching for something about someone, the key to our search is the person's name, as we would say it.  I don't think "What was Marx, Harpo's birth name?" (Adolph, changed to Arthur by 1911).  I think "What was Harpo Marx's birth name?" or just "What was Harpo's birth name?" (or I could just honk and whistle).  Either way, I start with "Harpo", and sure enough, "Harpo Marx" is on the list by the time I've typed that.  Other systems work similarly.

Searching by typing something in and expecting results to come back is quintessentially webby, and the autocomplete box is Web 2.0 in particular (whatever Web 2.0 is).

Friday, April 6, 2012

Old albums

A few weeks ago the Encyclopædia Britannica finally threw in the towel and, after over 200 years, stopped publishing the lovely multi-volume sets that have graced bookshelves the world over.  Naturally, there has been a run on the last edition (2010).  When I heard the story on the radio today there were supposed to have been no more than 800 copies left.  I'd be surprised if there were any left by now.

Clearly the whole point of this run is to get at the physical volumes.  The contents can be had digitally for much less.  The doorstop edition is valuable for the same reason any artifact is valuable apart from any utility it may have: rarity and emotional significance.

Imagine you are rummaging through an attic trying to decide what to keep and what to throw.  You run across a stash of vinyl LPs of popular hits from the 70s.  Odds are most if not all of the songs can be had digitally with better sound, but that's not the point.  As with the 2010 Britannica it's the physical artifact that matters.  Do you like that vintage artwork on the jacket with the circular imprint of the record worn into it?  Do you enjoy the tactile experience of dropping the needle on the platter, the crackle and pop of surface noise, the ritual of cleaning any wayward lint from the grooves?

Then you run across an album of photos, page after plain page of pictures tucked into little white corner-pockets, colors desaturated, edges curling.  Tucked into an envelope with them are the negatives.  Scan them and you probably have images of reasonable quality that you can attach to an email, share on your favorite social site and archive durably.  The physical artifact is less important here.  It's the actual images that matter, images you can't get anywhere else.  With the bits, you could create another album as good as the original one.

That's the common question that determines what's really of interest: what can't you get anywhere else?  It's not a matter of songs versus pictures or LPs versus photos.  If the vinyl in the Greatest Hits album is warped and cracked and the album art is nothing special, you may as well just buy the tunes online.  If the photo album is something your great Aunt put together, with cutouts and notes and decorations, you probably want the physical album as much as the images.

If the content is important, then you'll want to get it into the cloud, or at least into bits on some local disk.  If the artifact is important, then the web will play less of a role.

What's twice as big as the internet?

(Yikes ... I went 0 for March!)

I've mentioned before that telescopes can generate a lot of data.  IBM seems inclined to drive the point home by collaborating with ASTRON (the Netherlands Institute for Radio Astronomy) to put together "exascale" computing horsepower behind the world's largest radio telescope.

The telescope is actually (or rather, will be) an array of millions of antennas spread out over a square kilometer, from which the name SKA, for Square Kilometer Array.  This array is expected to produce on the order of an exabyte of data per day.  This is an absolutely ridiculous amount of data by today's standards.  Think one million terabyte disk drives, or twenty million feature film's worth of Blu-ray, or ... according to IBM, twice the daily volume currently carried on the internet.

I'm a little skeptical as to exactly how one measures that, but hey, you've got to trust a press release, right?


So where do you put an exabyte a day worth of data?  Well, you don't.  You're certainly not going to upload it to the web.  Particle physicists are faced with the same problem of having to figure out what portion of a huge data set to keep for later analysis, and a large part of running an experiment is setting up the "trigger" criteria by which the software collecting the data will decide what to keep and what to throw.  IBM and ASTRON's system will be dealing with the same problem, but on an even larger scale.

Or I suppose you could sign up two million people and somehow stream an equal share of the data to each at Blu-ray resolution all day every day, but somehow I doubt that kind of crowdsourcing will help much.

Friday, February 24, 2012

Is it OK to tweet "fire" in a crowded theater?

Evidently not.

Or at least, it's not a good idea to tweet in jest that you'll blow an airport sky-high if it remains closed for snow, so preventing you from visiting your girlfriend.  Paul Chambers of Doncaster, England found this out the hard way, paying a fine of £1000, gaining a criminal record and losing his job in the bargain.  His appeal will be heard before the high court of the UK and his defense has had at least one high-profile fundraiser, but it's all a bit sobering, to say the least.

This lack of humo(u)r on the part of airport security is not new, by the way, nor limited to the UK.  I remember as a kid -- so, ahem, well before 9/11 -- noticing a sign at the airport we were flying out of saying it was a federal crime even to joke about hijacking, bombs and such, and promptly blanching and making a mental note not to make any smart comments to the nice folks by the metal detector.

With that in mind, the remarkable aspect of the case isn't so much that it involves Twitter, though it is one of the first such cases, but that the authorities chose to prosecute for this particular remark at all.  I don't know how often such cases are prosecuted, but I'd guess it's not too often.  They certainly don't seem to make the press much.  I doubt the story would have been less remarkable had Mr. Chambers been brought in for making the same remark in person at the ticket counter.

In any case, caveat tweetor.

[Paul Chambers' conviction was eventually quashed, two and a half years later on the third appeal, the case having attracted considerable attention and celebrity involvement.  It's not clear if his job was reinstated, but according to Wikipedia he and his girlfriend did eventually marry.]

Saturday, February 11, 2012

Now I've seen everything

Sorry, horrible title.  I couldn't resist.

"Blind photographer" isn't a phrase that would spring to most people's minds readily, but not only are there such, there is -- of course -- a blog dedicated to blind photography.  From what I can tell, the photographers featured here aren't totally blind, but they are legally blind.  For example, I originally stumbled on this blog after reading about Craig Royal, who writes "My peripheral vision is blurred and the central vision is obscured by a white blindspot." and who processes his pictures with the aid of Photoshop and a telescope.

In other words, the blindness in question, while not complete, is very real and has a real effect on the images produced.  Indeed, the photographs on the site have a character all their own and, in my personal estimation, are just plain good art.

Clever copy and paste

Generally, if you select some part of a web page, copy it and paste it somewhere else, you'd expect to see pretty much what you'd selected, maybe with the formatting munged a bit.  Recently, though, I copied something (small enough for fair use) from one of the major sports outlets and was mildly surprised to see that it pasted with a handy "Read more" link including the URL of the article I'd quoted.  You can do that sort of thing in today's wonderful world of AJAX.

I suppose one could see this as an attempt to control copying of copyrighted material, which was muddled somewhere into my initial reaction, but really it seems like a more or less useful thing to do, and completely legitimate for a commercial publication.  For that matter, even in a non-commercial context attribution matters and an automatic backlink could be a nice feature.

Thursday, January 26, 2012

What, if anything, is a magazine?


A recent New York Times article tells the store of Esquire magazine's troubles in 2008 and 2009, and how it was able to survive them by adapting to the world of online publishing.

I'm not sure I've ever read Esquire in either print or digital form.  For that matter, I don't buy magazines much any more, but I do follow the (free) online content of some, particularly The Economist.

So do I, or does a digital Esquire subscriber, read magazines?  Pretty clearly yes, just as there's pretty clearly more to a magazine than its print edition.  So what's a magazine?  Some thoughts:
  • A classic magazine is almost always periodical, though a few publish irregularly.  On the web, content generally goes up when it's ready, regardless of the print publishing schedule.  Let's say a magazine is an ongoing publication.  There may or may not be a sequel to your favorite book, but part of publishing a magazine is the promise that there will be more.
  • A magazine is not tied to any particular individual.  Even in cases like Forbes or Oprah, where a particular individual's identity is an integral part of the brand, the actual magazine is the work of many people.  It is an institution, that can survive the departure of any particular person (though in some cases better than others).  This is where we can probably best see the tie to the earlier sense of magazine as a storehouse, and it's also a distinguishing feature between an online magazine and a blog.
  • Even though it's a group effort, a magazine does have a personality, or at least a good one does.  Even if its contributors don't always see eye to eye, there will be something about having that particular mix of opinions and styles that makes the magazine what it is.
From this point of view, as long as there are ongoing publications with multiple contributors and a recognizable personality, there will be magazines, regardless of the actual mechanics of publishing.

A corollary to that is that there ought to be just as much of a market for magazines as there ever was.  The puzzle, as always, is reaching that market and making sure everyone still gets paid, which is why I find it interesting that the headline of the Times article is in past tense: "How Esquire Survived ...", not "How will Esquire survive ..."

Wednesday, January 18, 2012

Does there have to be an app for that?

Weddings are generally public affairs, and they always have been.  I doubt it's ever been particularly difficult to find out who's planning a public wedding and when in a given area.  With the advent of online wedding planning it's now perhaps a bit easier yet, and if you're looking for a wedding to crash, well, there's an app for that.

Now, I'm with the author of the article in thinking that the kind of person who would use such a thing has -- how shall we say -- issues, but the flip side of that is, who's actually going to use it, as opposed to just having a laugh looking up one's friends and acquaintances?  Or more precisely, who's going to use it who wouldn't have been willing and able to crash a given wedding anyway?

In general, there's a lot of gray area when it comes to "enabling technologies", not to mention the larger sticky issue of to what extent technology can or should be considered without considering its potential consequences.  On the one hand, it's easy to say "The real problem is the wedding sites' privacy models.  The app just pulls together information that's already available."  But that's a cop-out.  As we've seen, pulling together information that's already available and making it universally accessible (if not useful) can make a significant difference.  Sometimes this is good, sometimes not, and just because something can be done doesn't mean it should be.

Just how much of a difference pulling together existing information and making it easy to get to can make depends on what the information is, how hidden it was, who wants to know and a host of other factors.  In this particular case, I doubt the app will make much difference.  That's not to condone wedding crashing, or the app, or to excuse its creators.  If your wedding is crashed by some tech-savvy boor who would otherwise have missed out, you have my sympathies, for whatever that's worth.

Saturday, December 31, 2011

Rumours and tweets of rumours

Someone at the Guardian (aided by academics at several universities) put in a bunch of overtime analyzing the Twitter traffic from last summer's riots in England.  In all, they traced seven rumors, five that turned out to be false, one that turned out to be true, and one they classify as "unsubstantiated."  They then put together a nice interactive graphic of the results, including a graph of the volume of traffic over time and a sort of cloud diagram color coded to show support for, opposition to, questioning of and commentary on the rumor in question, with size indicating the "influence" of the tweet, based on number of followers the originator of the tweet had.

The results are fascinating.  You should probably have a look at them yourself (here's the link again) before going on.


There is a fairly widespread notion that the web corrects itself.  People may put up misinformation, whether deliberately or in good faith, but eventually the real story will come out and supplant it.  The lead-in to the Guardian interactive graphic says so in as many words: "... Twitter is adept at correcting misinformation ..."

I don't see a lot of support for this in the data presented.

In the self-correcting model, you would expect to see an initial wave of green for a false rumor, coming with the original misinformation, steadily replaced by red, with possibly some yellow (questioning) and gray (commentary) in between.  Following is what actually happened for the five rumors determined to be definitely false.
  • Rioters attack London Zoo and release animals:  Initially, green traffic grows.  After a while, red traffic comes in denying the rumor.  Hours later, there is influential red traffic, but the green traffic is still about as influential.  Traffic then dwindles, with the last bits being green, still supporting the rumor hours after it has been disputed.
  • Rioters cook their own food in McDonalds: This one was picked up early by the website of the Daily Mail, which stated that there had been reports of this happening.  In any case, the green traffic surges moderately twice, before peaking at high volume several hours later.  There is no red traffic to speak of.
  • London Eye set on fire:  This one actually does follow the predicted pattern.  The initial green is quickly joined by yellow and red.  The proportion of red steadily grows, and as traffic dies down it is almost entirely red.
  • Rioters attack a children's hospital in Birmingham:  In this case one source of denials was someone actually working at the hospital.  Again, a strong surge of green is gradually taken over by red, but not completely.  As traffic dies down, the rumor is still being circulated as true.  Late in the game, it resurges again, though again there is a countersurge of denial.
  • Army deployed in Bank [I believe this refers to the area in London near the Bank of England and the Bank tube station]:  Traffic starts out yellow, as a question over a photo (which was actually a photo of tanks in Egypt).  Red traffic begins to grow, but so does green, and yellow continues to dominate.  Eventually everything dies down.  The last bits of traffic are yellow
In summary: One of the five cases follows the "good information drives out bad" model.  One other more or less follows it.  Two are an inconclusive mix of support and denial.  One consists almost entirely of support for a false rumor.

This was in one of the world's most connected cities, with widespread access to the internet, cell phones, land lines, television, newspapers, live webcams and whatever else.  Only in the case where the rumor was trivial to refute (for example via this webcam) did Twitter appear to self-correct. 

One would be hard-pressed, I think, to distinguish between the actual true rumor (Miss Selfridge set on fire -- that's the name of a store, not a person) and the false rumor about McDonalds based solely on the volume and influence of tweets confirming and denying.  Likewise, the unsubstantiated rumor (Police 'beat 16-year-old girl') follows its own pattern, mostly surges of green, but interspersed with yellow.

This may seem like a lot of argumentation just to say "Take your tweets with a grain of salt", but pretty much everything tastes better with data.