Weddings are generally public affairs, and they always have been. I doubt it's ever been particularly difficult to find out who's planning a public wedding and when in a given area. With the advent of online wedding planning it's now perhaps a bit easier yet, and if you're looking for a wedding to crash, well, there's an app for that.
Now, I'm with the author of the article in thinking that the kind of person who would use such a thing has -- how shall we say -- issues, but the flip side of that is, who's actually going to use it, as opposed to just having a laugh looking up one's friends and acquaintances? Or more precisely, who's going to use it who wouldn't have been willing and able to crash a given wedding anyway?
In general, there's a lot of gray area when it comes to "enabling technologies", not to mention the larger sticky issue of to what extent technology can or should be considered without considering its potential consequences. On the one hand, it's easy to say "The real problem is the wedding sites' privacy models. The app just pulls together information that's already available." But that's a cop-out. As we've seen, pulling together information that's already available and making it universally accessible (if not useful) can make a significant difference. Sometimes this is good, sometimes not, and just because something can be done doesn't mean it should be.
Just how much of a difference pulling together existing information and making it easy to get to can make depends on what the information is, how hidden it was, who wants to know and a host of other factors. In this particular case, I doubt the app will make much difference. That's not to condone wedding crashing, or the app, or to excuse its creators. If your wedding is crashed by some tech-savvy boor who would otherwise have missed out, you have my sympathies, for whatever that's worth.
Wednesday, January 18, 2012
Saturday, December 31, 2011
Rumours and tweets of rumours
Someone at the Guardian (aided by academics at several universities) put in a bunch of overtime analyzing the Twitter traffic from last summer's riots in England. In all, they traced seven rumors, five that turned out to be false, one that turned out to be true, and one they classify as "unsubstantiated." They then put together a nice interactive graphic of the results, including a graph of the volume of traffic over time and a sort of cloud diagram color coded to show support for, opposition to, questioning of and commentary on the rumor in question, with size indicating the "influence" of the tweet, based on number of followers the originator of the tweet had.
The results are fascinating. You should probably have a look at them yourself (here's the link again) before going on.
There is a fairly widespread notion that the web corrects itself. People may put up misinformation, whether deliberately or in good faith, but eventually the real story will come out and supplant it. The lead-in to the Guardian interactive graphic says so in as many words: "... Twitter is adept at correcting misinformation ..."
I don't see a lot of support for this in the data presented.
In the self-correcting model, you would expect to see an initial wave of green for a false rumor, coming with the original misinformation, steadily replaced by red, with possibly some yellow (questioning) and gray (commentary) in between. Following is what actually happened for the five rumors determined to be definitely false.
The results are fascinating. You should probably have a look at them yourself (here's the link again) before going on.
There is a fairly widespread notion that the web corrects itself. People may put up misinformation, whether deliberately or in good faith, but eventually the real story will come out and supplant it. The lead-in to the Guardian interactive graphic says so in as many words: "... Twitter is adept at correcting misinformation ..."
I don't see a lot of support for this in the data presented.
In the self-correcting model, you would expect to see an initial wave of green for a false rumor, coming with the original misinformation, steadily replaced by red, with possibly some yellow (questioning) and gray (commentary) in between. Following is what actually happened for the five rumors determined to be definitely false.
- Rioters attack London Zoo and release animals: Initially, green traffic grows. After a while, red traffic comes in denying the rumor. Hours later, there is influential red traffic, but the green traffic is still about as influential. Traffic then dwindles, with the last bits being green, still supporting the rumor hours after it has been disputed.
- Rioters cook their own food in McDonalds: This one was picked up early by the website of the Daily Mail, which stated that there had been reports of this happening. In any case, the green traffic surges moderately twice, before peaking at high volume several hours later. There is no red traffic to speak of.
- London Eye set on fire: This one actually does follow the predicted pattern. The initial green is quickly joined by yellow and red. The proportion of red steadily grows, and as traffic dies down it is almost entirely red.
- Rioters attack a children's hospital in Birmingham: In this case one source of denials was someone actually working at the hospital. Again, a strong surge of green is gradually taken over by red, but not completely. As traffic dies down, the rumor is still being circulated as true. Late in the game, it resurges again, though again there is a countersurge of denial.
- Army deployed in Bank [I believe this refers to the area in London near the Bank of England and the Bank tube station]: Traffic starts out yellow, as a question over a photo (which was actually a photo of tanks in Egypt). Red traffic begins to grow, but so does green, and yellow continues to dominate. Eventually everything dies down. The last bits of traffic are yellow
In summary: One of the five cases follows the "good information drives out bad" model. One other more or less follows it. Two are an inconclusive mix of support and denial. One consists almost entirely of support for a false rumor.
This was in one of the world's most connected cities, with widespread access to the internet, cell phones, land lines, television, newspapers, live webcams and whatever else. Only in the case where the rumor was trivial to refute (for example via this webcam) did Twitter appear to self-correct.
One would be hard-pressed, I think, to distinguish between the actual true rumor (Miss Selfridge set on fire -- that's the name of a store, not a person) and the false rumor about McDonalds based solely on the volume and influence of tweets confirming and denying. Likewise, the unsubstantiated rumor (Police 'beat 16-year-old girl') follows its own pattern, mostly surges of green, but interspersed with yellow.
This may seem like a lot of argumentation just to say "Take your tweets with a grain of salt", but pretty much everything tastes better with data.
Monday, November 28, 2011
Voices from the dashboard
All my life I've taken road trips, partly by natural inclination, partly by necessity. It's a largely timeless experience. Sure, the roads have improved (see the Grapevine Grade section of this page for a good example), the speed limits are higher, cars are faster and safer and there's not a lot of "local flavor" in most stopping points unless you actively seek it out, but for the most part road trips have been road trips since well before Kerouac.
One thing that has changed is the soundtrack, and not just because tastes in music have changed. When I was a kid, any audio not provided by the car and its occupants came from the radio, and if you were on a long haul, it was the AM radio. Keeping FM tuned in was and remains too much of a hassle. An AM station, especially one of the "clear channel" stations (not to be confused with the media conglomerate) licensed to broadcast at high power, could be good for hours -- enough for a whole sports fixture, several runs through the news or all the whacked-out talk radio conspiracy theories you could eat.
The key feature here, particularly on a solo trip through, say, the desert southwest US, was the lack of choice. You'd be doing well to have your pick of baseball, UFO speculation and the company of your own thoughts, and a hundred miles or so out of Albuquerque on a dark night with the game a blowout the UFO speculation starts sounding interesting and plausible.
By the time I was doing my own solo long hauls, cassette tape was an option, but a library of a few dozen albums can be limiting after a while -- and suppose you want to know what's going on in the world, or just let someone else handle the programming for a while? The in-dash CD (briefly supplemented by a multi-disc changer in the trunk) increased one's options, but the same basic constraints applied. Only with the advent of satellite radio was there little reason to tune in to local stations at all.
And now there's the web. As long as you've got a smartphone, bars, a bit of cable and an aux input, you can listen to pretty much anything. Stream your favorite home station. Stream your favorite internet station. Play your podcasts. Dial up Pandora. AM won't be completely disappearing anytime soon -- technologies written off as obsolete seldom do -- but the proportion of people who know or care must be steadily dwindling. Likewise I'd rather not try to predict whether or when web audio will supplant satellite radio, but if I had to place long-term bets, I'd bet on the web.
It's hard to argue that having a huge palette of choices isn't progress of some sort, but there's something to be said for being drawn out of one's comfort zone because there's only one game in town.
One thing that has changed is the soundtrack, and not just because tastes in music have changed. When I was a kid, any audio not provided by the car and its occupants came from the radio, and if you were on a long haul, it was the AM radio. Keeping FM tuned in was and remains too much of a hassle. An AM station, especially one of the "clear channel" stations (not to be confused with the media conglomerate) licensed to broadcast at high power, could be good for hours -- enough for a whole sports fixture, several runs through the news or all the whacked-out talk radio conspiracy theories you could eat.
The key feature here, particularly on a solo trip through, say, the desert southwest US, was the lack of choice. You'd be doing well to have your pick of baseball, UFO speculation and the company of your own thoughts, and a hundred miles or so out of Albuquerque on a dark night with the game a blowout the UFO speculation starts sounding interesting and plausible.
By the time I was doing my own solo long hauls, cassette tape was an option, but a library of a few dozen albums can be limiting after a while -- and suppose you want to know what's going on in the world, or just let someone else handle the programming for a while? The in-dash CD (briefly supplemented by a multi-disc changer in the trunk) increased one's options, but the same basic constraints applied. Only with the advent of satellite radio was there little reason to tune in to local stations at all.
And now there's the web. As long as you've got a smartphone, bars, a bit of cable and an aux input, you can listen to pretty much anything. Stream your favorite home station. Stream your favorite internet station. Play your podcasts. Dial up Pandora. AM won't be completely disappearing anytime soon -- technologies written off as obsolete seldom do -- but the proportion of people who know or care must be steadily dwindling. Likewise I'd rather not try to predict whether or when web audio will supplant satellite radio, but if I had to place long-term bets, I'd bet on the web.
It's hard to argue that having a huge palette of choices isn't progress of some sort, but there's something to be said for being drawn out of one's comfort zone because there's only one game in town.
Wednesday, November 23, 2011
In which a theorist discovers something unsettling, exhilarating or both
There seems to be a natural human compulsion to keep checking the soup to see if it's boiling, to check the weather, to check the latest sports scores and stock prices, to check for messages, and so on and so forth. One of the less savory properties of the web is that it provides the means to indulge this compulsion to the nth degree.
I personally try to steer clear of this, which is the main reason I'm not on Facebook or Twitter (and not particularly active on Google+), but I'm certainly not immune. Are there any comments on Field Notes? Has anyone read the latest brilliant post (there are at least three ways to check, each giving its own opinion)? Anything new on the few sites I do follow?
Since I'm not on Facebook, I don't play Facebook games, but evidently a lot of people do. Zynga's Farmville, for example, has over 80 million subscribers, still a small minority of the gazillion on Facebook, but a big number in most normal contexts. This has irked traditional computer game creators, sucked up untold hours of human life, and intrigued computer gaming analyst/critic Ian Bogost.
Bogost noted that games like Farmville involve relatively little actual gameplay. Rather, it's the social aspect that seems to dominate. This is nothing new in gaming, but again the natural "I need to check what's going on" factor of the web in general and Facebook in particular acts to intensify this. Bogost coined the term "Cow Clicker" to describe games like Farmville where the action seems to consist mainly in, for example, clicking on depictions of animals when various timers run out.
Unable to leave it at that, Bogost took the next logical step and created a Facebook game called Cow Clicker designed to distill the social gaming experience to its purest elements. It goes like this:
- You have a picture of a cow on your page.
- You click on it.
- It does nearly nothing -- I think maybe it moos or otherwise makes a sound?
- You can't click again for six hours.
Yep. That's my story and I'm sticking to it.
If you don't want to wait six hours, you could spend "mooney" -- Cow Clicker's own virtual currency -- to get the right to click sooner. You could earn mooney by clicking on your cow, by having your friends click on feed stories about you clicking on your cow, or by paying a small amount of actual money.
People played this. Not 80 million, but somewhere around 50,000, not too bad for a joke of a game with no marketing behind it.
Clearly the actual cow clicking is a MacGuffin. No one cares much about it. What people care about is whether their friends are also playing and clicking on their feed stories, thereby generating not just more mooney, but, crucially, another thing to check in on.
Bogost had mixed feelings about this. Among other things, he found himself, despite his intentions, checking in on whether people were playing the game and what they wanted from it.
Naturally, people wanted upgrades. They wanted their choice in cows. Cowthulhu was a popular request. Eventually Bogost put up an "app store" with a selection of cows, and (I gather) added another feature or two. If you were really hardcore, you could pay $100 (or the equivalent in mooney from whatever source) for Bling Cow. Why on earth would anyone do this? Well, your friends would all know that you had splashed out for the Bling, and wouldn't they be envious? Again, people actually did this.
Eventually, Bogost was unable to shake the feeling he'd created a monster, and so he brought about the Cowpocalypse. At a preset time -- which players would hasten by actually playing the game but could defer by, yep, paying mooney -- the cattle would all be "raptured", leaving only the empty spaces on which they had once stood. And so the Cowpocalypse eventually came to pass.
At this point, it may not come as a shock that people kept playing. To recap: people were now paying (small amounts of) money for the privilege of clicking on an empty space and letting their friends know about it.
You couldn't ask for a better illustration that when economists talk about "rational consumers", they only mean people that behave as though there's some sort of "utility function", be it ever so screwy, that they're bent on maximizing. "Rational" in the usual sense has got nothing to do with it.
If people were actually rational in the usual sense, Cow Clicker would never have happened, but of course they aren't. We are, at a very basic level, social animals. We want to know what other people are doing. What in particular they're actually doing is often much less important to us than whom they're doing it with and the fact that we know this. If the entirety of Facebook were pushing a button from time to time saying "I'm here", selecting people to notify of that and having the system tell people you're notifying know whom else you're notifying, it would not be outlandish to think people would still use it.
The cynic would say that that really is the essence of Facebook and "social networking" in general, but I wouldn't go quite that far. I said above that what people are doing is often much less important than knowing it and knowing who knows, but that doesn't mean it's always more important. Content can matter -- of course -- but it's worth noting that it doesn't always.
Labels:
common knowledge,
facebook,
game theory,
Ian Bogost,
social networks
Monday, November 7, 2011
Yay! Yet another way to spam!
While buying something online today, I was presented with a popup asking me if I wanted to chat live with a representative about what looked like a loyalty program. I went ahead and clicked, even though my spidey-sense told me not to.
PhineasTaylor is typing ...
Hello there! Thank you for taking a moment to chat with me about the wonderful opportunity of joining buyeverythingthroughus.com. With buyeverythingthroughus.com, etc., etc.OK, a little boilerplate to get things going. Hang on though, there's more
PhineasTaylor is typing ...
Buyeverythingthroughus.com will improve your life in every possible way. It will make you rich and famous. It will cure dandruff and halitosis. Children will love you. Adults will want to be you. Your friends will adore you. Your enemies will envy you and then slink away in shame and fear, etc., etc.Right ... anything else?
PhineasTaylor is typing ...This P.T. person sure types a lot.
Buyeverythingthroughus.com will cure hunger. It will bring about world peace and universal prosperity. Yankees and Red Sox fans will embrace each other with love in their eyes [well, maybe it didn't go quite that far].Since this is ostensibly a person typing, it's coming across slowly enough that there's plenty of time to go googling and find out that buyeverythingthroughus.com is about what you'd expect it is.
Knowing all that, what do you say to this exciting opportunity?I said "No, thank you" and dismissed the chat window. I couldn't help wondering, though, whether whoever coded this up had the chutzpah to submit a paper on an exciting new "intelligent agent".
Banking on web security
People do care about web security. There are highly competent full-time professionals in the field. There are conferences on the subject on a regular basis. You'll see them in the press -- Experts Meet to Fix Security on the Web.
And yet, in large part because the problems to be solved are hard and involve significant non-techical factors, there is no shortage of things that could stand to be fixed.
And yet, in large part because the problems to be solved are hard and involve significant non-techical factors, there is no shortage of things that could stand to be fixed.
- Authentication is a mess. For the most part, we have passwords and security questions. I've griped about this before, multiple times, and I'm sure I'll gripe about it again.
- Identity is a mess. Everyone has scads and scads of identities -- logins here, there and everywhere. They can easily get confused ("That wasn't me, that was some other David Hull!"). There's no good way to say two random identities are or aren't the same. I've griped and speculated about this before, too, and I expect I'll have more to say on that, too.
- Anonymity is problematic. Everything you do on the web leaves traces, but unless you're paying extremely close attention you generally don't know exactly what kind, or whether they can be tied to your identity (whatever that is).
- Network infrastructure is scary. Https with certificates is widely deployed, and most people probably at least know that some sites are "secured" and some aren't, but many fewer understand (or should need to understand) details like signatures, secure hashes and certificate authorities, or what can fail and what's less likely to. Did I mention DNS?
- PCs are scary. Viruses, rootkits, system crashes ... some platforms are better designed than others, but nothing's perfect.
- The cloud has its own problems. Who owns what you put there? Who's liable if data is lost or compromised? Who can see what? Who can see who sees what?
- Spam is a perennial problem, not helped by any of the above.
I could go on, but if it's so bad -- and it is -- how does it work at all? People continue to be able to use credit cards both online and in person, people continue to email and text each other all sorts of sensitive information, people continue to turn to the web for all sorts of vital information. Clearly Bad Things can happen to a person on the web, but just as clearly it's not bad enough often enough to put people off the web entirely. Far from it.
My guess is that banks have a lot to do with it, at least in the US. In particular
- Banks handle liability. If someone steals your credit or debit card, whether physically or online, you can tell your bank and generally they will make sure you don't have to pay for things you didn't buy. That's oversimplified, and there are certainly cases where that simple process has turned into a nightmare, but it's still a vital part of getting people to do business confidently online.
- Bank cards provide a de facto stable identity. If you're buying something from my web site, I do care who you are (well, I would, and stores in general do seem to care what their customers are up to), but I certainly also care that your payment is going to go through. To some extent I'm talking to you, but I'm also talking to your bank account.
On the first point, you're not responsible for keeping your bank accounts absolutely safe. You're responsible for taking reasonable precautions, so that if someone does get hold of your account number and misuses it, they're clearly at fault (the usual "I'm not a lawyer" disclaimer applies here). Putting the rest of the burden on the banks and legal system is a large part of what keeps the wheels turning.
On the second point, if I shop at store A and store B, it's important that my bank knows that those purchases both come out of my account, and I know that I'm the same person in both cases (at least on a good day). It's less important that store A and store B know I'm the same person. There may even be cases where I'd rather they didn't know.
In short, security and identity matter when money is at stake, in which case your accounts serve as your identity and you have legal protections that predate the web.
Security and identity also matter where reputation is at stake, that is in the social realm, be it email, social networks, Twitter or whatever. The landscape is different there, but it's worth noting that most accounts and identities, including your bank accounts, don't play into that much. If someone compromises my account at widgetco.com, they might be able to have a truckload of widgets sent to my address at my expense, but they won't be able to say embarrassing things about me on this blog. Likewise if they compromise my bank account, though that would of course be bad for other reasons.
If you buy that, then you should make sure to use strong unique passwords and unique security questions for your bank accounts, your email accounts and your major social accounts, and use better security than that when it's available. How much to worry about other accounts depends on how closely they're tied to the accounts that matter. For example, if your city's online parking ticket paying site doesn't remember credit card numbers or your nefarious history of overparking, you probably don't care as much about security there.
Labels:
anonymity,
e-commerce,
law,
passwords,
security
Friday, October 14, 2011
Dennis Ritchie, 1941 - 2011
I have no intention of turning this blog into an obituaries column, and no desire to see "celebrity deaths come in threes" spill over into the tech world, but having noted the passing of Steve Jobs I feel obliged to note the passing of Dennis Ritchie as well.
You may or may not have heard of him before. It took the major news outlets a while to pick up the story, and even then it wasn't front page. For hours the main public source was colleague Rob Pike's Google+ page. That's not too surprising. CEO of major corporations and eminent computer scientist are two completely different gigs. Nonetheless, Ritchie had as profound an effect on the Web As We Know It as anyone else, even though his groundbreaking work predates the web by a good measure.
It's fair to say that the web as we know it would not exist if not for Unix. The first web server ran on NeXTSTEP, which traces its roots to Unix [and, in fact, NeXT was run by the late Steve Jobs -- tech is a small world at times -- D.H. Nov 2018]. A huge number of present-day web servers, large and small, run on Linux/GNU which, even though the Linux kernel was developed from scratch and GNU stands for "GNU's Not Unix", provide an environment that's firmly in the Unix lineage. The HTTP protocol the web runs on has its roots in the older internet protocols and belongs to a school of development in which Unix played a major role.
Ritchie was one of the original developers of Unix.
The Unix operating system, the Linux kernel, many of the GNU tools and countless other useful things (and at least one lame hack) are written in the C language, which is also one of the bases for C++, C#, Objective C and Java, among others. All in all, C and its descendants account for a large chunk of the software that makes the web run, and for years, before the ANSI C standard, the de facto standard for the language was a book universally called "K&R" after its authors, Brian Kernighan and Dennis Ritchie. That flavor of the language is still called "K&R C".
Ritchie continued to do significant work throughout his life and won various high honors, including the Association for Computing Machinery's top honor, the Turing award, and the US National Medal of Technology. He was head of the Lucent Technologies System Software Research Department when he retired in 2007. He may not have been a cultural icon, but in the world of software geekery he cast a long shadow.
RIP
Thursday, October 6, 2011
So ... what version are we on?
Trying to do a bit of tidying up, I tagged a previously-untagged recent post "Web 2.0". I did this because the post was a followup to an older post that was specifically about Web 2.0, but it felt funny. Web 2.0 is starting to sound like "Information Superhighway" and "Cyberspace". A quick check of the Google search timeline for the term suggests that usage peaked around 2007 and has been declining steadily since. Always on the cutting edge, Field Notes uses the tag most heavily in 2008.
Google's timeline isn't foolproof. Anything given a date before the late 90s is probably an article that mentioned the date (and Web 2.0) and gave no stronger indication of when the page is from. On the other hand, the more recent portion is probably more representative, since there's more metadata around these days. Also, the numbers are larger, which is often good for washing out errors.
But anyway, are we still in Web 2.0? Are we up to 3.0? Does it really matter (spoiler: probably not)?
I've argued before that while Web 1.0 was a game-changing event, Web 2.0 is more a collection of incremental improvements. Enough incremental improvements can produce significant changes as well, but not in such a way as you can draw a clear bright line between "then" and "now". The Linux kernel famously spent about 15 years on version 2.x, only just recently moving up to 3.0, and Linus says very clearly that 3.0 essentially just another release with a shiny new number. From a technical standpoint I'd say we've been on Web 2.x for a while and will continue to be for a while, unless we decide to start calling it 3.x instead.
Because, of course, "Web 2.0" is not a technical term. Never mind who uses it to what ends in what context. The ".0" gives the game away to begin with. A real version 2.0, if it ever exists, is very soon supplanted by 2.0.1, or 2.1, or 2.0b or whatever as the inevitable patches get pushed out, which is why I was careful to say "2.x" above. "2.0" as popularly used doesn't designate a particular version. It's supposed to indicate a dramatic change from crufty old 1.0 (or 1.x if you prefer). In the real world of incremental changes, that trope will only get you so far.
Hmm ... in real life versioning usually goes more like
Google's timeline isn't foolproof. Anything given a date before the late 90s is probably an article that mentioned the date (and Web 2.0) and gave no stronger indication of when the page is from. On the other hand, the more recent portion is probably more representative, since there's more metadata around these days. Also, the numbers are larger, which is often good for washing out errors.
But anyway, are we still in Web 2.0? Are we up to 3.0? Does it really matter (spoiler: probably not)?
I've argued before that while Web 1.0 was a game-changing event, Web 2.0 is more a collection of incremental improvements. Enough incremental improvements can produce significant changes as well, but not in such a way as you can draw a clear bright line between "then" and "now". The Linux kernel famously spent about 15 years on version 2.x, only just recently moving up to 3.0, and Linus says very clearly that 3.0 essentially just another release with a shiny new number. From a technical standpoint I'd say we've been on Web 2.x for a while and will continue to be for a while, unless we decide to start calling it 3.x instead.
Because, of course, "Web 2.0" is not a technical term. Never mind who uses it to what ends in what context. The ".0" gives the game away to begin with. A real version 2.0, if it ever exists, is very soon supplanted by 2.0.1, or 2.1, or 2.0b or whatever as the inevitable patches get pushed out, which is why I was careful to say "2.x" above. "2.0" as popularly used doesn't designate a particular version. It's supposed to indicate a dramatic change from crufty old 1.0 (or 1.x if you prefer). In the real world of incremental changes, that trope will only get you so far.
Hmm ... in real life versioning usually goes more like
- 0.1, 0.2 ... 0.13 ... 0.42 ... 0.613 as we sneak in "just one more" minor tweak before officially turning the thing loose
- 1.0 First official release. Everyone collapses in a heap. The bug reports start coming in
- 1.1 Yeah, that oughta fix it.
- 1.1.1, 1.1.2 ... 1.1.73 ... the third number emphasizing these are just "small patches" to our mostly-perfect product -- bug fixes, cosmetic changes, behind-the-scenes total rewrites, major new features important customers were demanding, that sort of thing.
- 2.0.1 OK, now we've got some snazzy new stuff. Anything coming up for a while is just going to be a "minor update". Everyone collapses in a heap. Bug reports keep coming in.
- 2.0.2, 2.0.3 ... yeah, we've seen this movie before
- 5.0, because our latest version is so much better than anything you've ever seen, including our own previous versions (Actually, version 3.x ended in tears, 5.x is largely a rewrite by a different team and no one knows what happened to 4.x -- maybe that's why one of the co-founders was sleeping under their desk and living on pizza for a couple of months?).
- 5.0.1, 5.0.2 ... you know the drill
- Artichoke. Yep. Artichoke. Version numbers are so two-thousand-and-late [already well out of date when I wrote that ... how meta ... -- D.H Dec 2018]. We're going with vegetables now. Already having long meetings on whether it's Brussels Sprout or Broccoli next.
- Artichoke 1.1, Artichoke 1.2 ...
Wednesday, October 5, 2011
Steve Jobs, 1955-2011
Well, we all knew it was coming, but you could still feel the earth shift. None of us in the tech business has remained untouched by Jobs' work, and by extension, Jobs himself. There was never, nor will there ever be, anyone quite like him.
RIP
RIP
Crowdsourcing the sky
Astronomy has been likened to watching a baseball game through a soda straw. For example, the Hubble Deep Field, assembled from 342 images taken over the course of ten days, covers about 1/500,000th of the sky, or about the size of a tennis ball seen a hundred yards away. It's quite possible to survey large portions of the sky, but there are trade-offs involved since you can only collect so much light so fast. To cover a large area and still pick up faint objects, you need some combination of a big telescope and a lot of time. The bigger the telescope (technically, there's more to it than sheer size) the faster you can cover a given area down to a given magnitude (how astronomers measure faintness).
The Large Synoptic Survey Telescope (LSST) is designed to cover the entire sky visible from its location every three days, using a 3.2 gigapixel camera and three very large mirrors. In doing this, it will produce stupefying amounts of data -- somewhere around 100 petabytes, or 100,000 terabytes, over the course of its survey. So imagine 100,000 terabyte disk drives, or over 2 million two-sided Blu-ray disks. Mind, the thing hasn't been built yet, but two of its three mirrors have been cast, which is a reasonable indication people are serious. Even if it's never finished, there are other sky surveys in progress, for example the Palomar Transient Factory.
Got a snazzy 100 gigabit ethernet connection? Great! You can transfer the whole dataset in a season -- start at the spring equinox and you'll be done by the summer solstice. The rest of us would have to wait a little longer. My not-particularly-impressive "broadband" connection gets more like 10 megabits, order-of-magnitude, so that'd be more like 2500 years, assuming I don't upgrade in the meantime and leaving aside the small question of where I'd put it all.
Nonetheless, the LSST's mammoth dataset is well within reach of crowdsourcing, even as we know it today:
Wikipedia references a 2007 press release saying Google has signed up to help. As usual I don't know anything beyond that, but it does seem like a googley thing to do.
The Large Synoptic Survey Telescope (LSST) is designed to cover the entire sky visible from its location every three days, using a 3.2 gigapixel camera and three very large mirrors. In doing this, it will produce stupefying amounts of data -- somewhere around 100 petabytes, or 100,000 terabytes, over the course of its survey. So imagine 100,000 terabyte disk drives, or over 2 million two-sided Blu-ray disks. Mind, the thing hasn't been built yet, but two of its three mirrors have been cast, which is a reasonable indication people are serious. Even if it's never finished, there are other sky surveys in progress, for example the Palomar Transient Factory.
Got a snazzy 100 gigabit ethernet connection? Great! You can transfer the whole dataset in a season -- start at the spring equinox and you'll be done by the summer solstice. The rest of us would have to wait a little longer. My not-particularly-impressive "broadband" connection gets more like 10 megabits, order-of-magnitude, so that'd be more like 2500 years, assuming I don't upgrade in the meantime and leaving aside the small question of where I'd put it all.
Nonetheless, the LSST's mammoth dataset is well within reach of crowdsourcing, even as we know it today:
- Galaxy Zoo claims that 250,000 people have participated in the project. Many of them are deadbeats like me who haven't logged in for ages, but suppose there are even 10,000 active participants.
- The LSST is intended to produce its data over ten years, for an average of around 2-3Gbps. Still fairly mind-bending -- about a thousand channels worth of HD video, but ...
- Divide that by our hypothetical 10,000 crowdsourcers and you get 200-300Kbps, not too much at all these days. Each crowdsourcer could download a 3GB chunk of data in under an hour in the middle of the night or spread it out through the day without noticeably hurting performance.
- Assuming you kept all the data, you'd need a new terabyte disk every few months, so that's not prohibitive either.
- The hard part is probably uploading a steady stream of 2-3Gbps (bittorrent wouldn't help here, since each recipient gets a unique chunk of data). As far as I can tell the bandwidth is there, but at that volume I'm guessing the cost would be significant.
- In reality, there would probably be various reasons not to ship out all the raw data in real time, but instead send a selection or a condensed version.
Wikipedia references a 2007 press release saying Google has signed up to help. As usual I don't know anything beyond that, but it does seem like a googley thing to do.
Labels:
astronomy,
crowdsourcing,
Galaxy Zoo,
Google,
ridiculous amounts of data
Monday, September 26, 2011
Real science, hot off the web
A while ago I commented on an Economist article claiming that Web 2.0 tools "were beginning to change the shape of the scientific debate." My contention was that the web wasn't so much changing the debate as changing the means of publication. In particular, there had always been a trade-off between speed of publication and thoroughness of review, and the web was becoming a publishing mechanism of choice on the lightly-reviewed end of that continuum.
More recently, looking for something I no longer recall, I ran across Cornell's arxiv.org (I assume the x is meant to represent a Greek χ), a repository for "Open access to 703,281 e-prints in Physics, Mathematics, Computer Science, Quantitative Biology, Quantitative Finance and Statistics." The number 703,281 was current when I scraped it. It's probably higher by now.
That's a lot of articles, but does anyone use it for anything important? Well, one recent entry is Superluminal neutrinos in long baseline experiments and SN1987a (Cacciapaglia, Deandrea, Panizzi et. al., yes, those neutrinos). Indeed, there seems to be a lot of activity in the experimental high-energy physics section overall, which makes sense. It's useful to have experimental results available quickly, bearing in mind that there can be quite a bit of calibration, number-crunching and checking before an experimental paper is ready for public consumption (months, in the case of the neutrino paper).
Submissions to arxiv.org "must conform to Cornell University academic standards". It's not immediately clear to me what process is in place to ensure this, but from a little browsing it's clear that these are serious academic papers. It also seems reasonable to assume that most of the papers have not been through the full process of peer review required for publishing in a major journal. Indeed, the published version is almost certainly not going to appear on such a site, if only for reasons of copyright.
It seems like a good niche to fill. If you have a significant result that you're comfortable sharing with the world and staking your reputation on, there should be some way to make it immediately available, with the implied tradeoff between speed and thoroughly careful vetting. Publishing under the aegis of a major university gives everyone some assurance that you're at least doing real research. I notice from random sampling that the reference sections generally don't cite arxiv.org, giving some indication of the preliminary nature of the publications.
With that in mind, arxiv.org looks like a great resource not only for working academics but for the curious layperson as well.
More recently, looking for something I no longer recall, I ran across Cornell's arxiv.org (I assume the x is meant to represent a Greek χ), a repository for "Open access to 703,281 e-prints in Physics, Mathematics, Computer Science, Quantitative Biology, Quantitative Finance and Statistics." The number 703,281 was current when I scraped it. It's probably higher by now.
That's a lot of articles, but does anyone use it for anything important? Well, one recent entry is Superluminal neutrinos in long baseline experiments and SN1987a (Cacciapaglia, Deandrea, Panizzi et. al., yes, those neutrinos). Indeed, there seems to be a lot of activity in the experimental high-energy physics section overall, which makes sense. It's useful to have experimental results available quickly, bearing in mind that there can be quite a bit of calibration, number-crunching and checking before an experimental paper is ready for public consumption (months, in the case of the neutrino paper).
Submissions to arxiv.org "must conform to Cornell University academic standards". It's not immediately clear to me what process is in place to ensure this, but from a little browsing it's clear that these are serious academic papers. It also seems reasonable to assume that most of the papers have not been through the full process of peer review required for publishing in a major journal. Indeed, the published version is almost certainly not going to appear on such a site, if only for reasons of copyright.
It seems like a good niche to fill. If you have a significant result that you're comfortable sharing with the world and staking your reputation on, there should be some way to make it immediately available, with the implied tradeoff between speed and thoroughly careful vetting. Publishing under the aegis of a major university gives everyone some assurance that you're at least doing real research. I notice from random sampling that the reference sections generally don't cite arxiv.org, giving some indication of the preliminary nature of the publications.
With that in mind, arxiv.org looks like a great resource not only for working academics but for the curious layperson as well.
Subscribe to:
Posts (Atom)
