Showing posts with label augmented reality. Show all posts
Showing posts with label augmented reality. Show all posts

Thursday, August 06, 2009

Link: An Illustrated Brief History of Augmented Reality

Check it out.

Plus who knew that Philippe Kahn invented the camera phone?

Apparently, a lot of people, just not me. The guy is so awesome.

Tuesday, July 07, 2009

Can iPhone et al Drag Augmented Reality Into Non-Augmented Reality?

I’ve been pitching augmented reality apps in startup circles for a few years now, so it was exciting to see the AR startup crop grab some press coverage this week (VentureBeat and more VentureBeat and …even SF ABC 7).

With the iPhone hardware suite (and comparable devices like the Android and Pre), there’s no shortage of hardware for the core AR tasks:

  • capture decent-resolution images
  • recognize “target” areas in images
  • contextualize the targets if necessary, by adding GPS data, solid-state compass data, and/or accelerometer (angle) data
  • lookup augmentation data suitable to the target and the end-user via suitable web services
  • employ “billboarding” or 3D rendering to composite a representation of the augmentation data on top of the target
  • repeat as fast as possible without draining the battery (yeah, right)

Now for part two of the plan: this facility needs to run through a cool looking visor (a.k.a. "head-mounted display” or HMD). And neither the $6,000+ price tag nor the Silence-of-the-Lambs-night-vision look is appealing on these traditional high-tech units.

Happily, there are mass-market headsets designed for the gaming or personal entertainment market which are almost ready to go. A couple are even within striking distance of cool factor. Maybe an Apple logo would be enough to do it, at least for the Bay Area.

Even better, leaders such as Vuzix recognize the need to provide video and accelerometer data from the POV of the headset (vastly reducing the amount of computation needed contextualize the image). They appear to be planning these capabilities as optional clip-on modules to their newest “Fall ‘09” model visor.

Note the word “planning.”

Like smartphones themselves, we’ve been here before … a lot of times. The iPhone was easily the industry’s 10th attempt at a commodity handheld computer, so it’s not like the writing is on the wall. Unless it’s AR writing:


Monday, May 26, 2008

My Brain on Web 3.0: a Killer App for the Semantic Web?

Here's a semantic app I'd really like and which could make the Internet -- and data stores in general -- more valuable. If anyone sees this and thinks it's a great startup idea, you're welcome to develop it.

Right now, a number of semantic analysis engines are being developed, and many are running well in production. Examples include ClearForest, which is related to Reuters Calais.

There are also a bunch of up-and-coming semantic-web apps, like Twine, that add semantic analysis to the extant web 2.0 experience. But while the RDF-enabled-del.icio.us-on-steroids-with-autotagging may be nice -- heck, I'm sure I'll be using one of those systems -- I want something that can carry out the kind of mental associations that I normally would have to do myself.

In order to make that a reality, I'll need several things:

  1. The Model: it works by association, and associations have direction, degree, and kind, among other things. So we need more than just a network. We need a model that implements a metric space or vector space, allowing distances to be computed between any two points with sensible behavior, and where measures (length, volume) can easily and intuitively work.
  2. Concepts (e.g., "politics"), and the places concepts are referenced (say, a political blog page), both live inside this space as subspaces, just like in my brain. Some of these spaces may consist of just one element.
  3. The formulae that support metric need to be fine tuned so that abstract tags ("politics", "justice" as opposed to "Davis, CA" or "bananas") don't suck a million other things to within epsilon of themselves. We can't have everything that linguistically has to do with politics cluster tightly in the space to a politics node, or else the system isn't terribly helpful.
  4. I want the system to start out thinking like me ... and then I can experiment later with "social thinking." What does this mean? My semantic tagging is different from everyone else's. Each group or culture I'm in sees the content differently from other groups and cultures; there is no universal invariant conceptual structure. One persons sees a news story and thinks "economics" while someone else thinks "environment" and another thinks "social justice." If we mush all these tags together we get nothing terribly useful. So: let's start out with my view of the world, we'll compare and integrate others' later.
  5. How to do #4? Start with every web page I visit (not just those I actively tag) -- read my history file or my network traffic (obviously, keep raw data local for now). Read my email and my calendar and my notes and phone (PIM) and my to-do list. And weight accordingly: associations in a web page I write (like this blog post) count more than stuff I browse through; notes I make in my phone or Outlook count for even more; the metadata in my calendar and the titles of my contacts mean a heck of a lot for the model. Walk my social graph and look at what my friends know and are interested in! These are all straightforward algorithmic steps. Leaving aside any self-tuning in the metrics engine, there is no AI or black box here.
  6. The data from #5 is part of the metric function ... that's how the system shapes itself to my view of the world, or at least my "attention waveform" as I transmit that through keystrokes and mouse clicks. Concretely, nodes "move" in the space based on whether I actually key them, whether they appear in meetings in my calendar, whether they are tightly clustered to my personal contacts etc. Even my contacts are arranged based on how long I've known them, what I talk to them about, how often, etc.
  7. The system will make mistakes. So all the more reason for (1) privacy around my core data and (2) a dashboard where I can "juice" certain things or move them around. (This is where interesting goal-oriented self-retuning can come inThere is plenty of other data that can be shared and deduced for network-effect-dependent revenue streams.
  8. A UI into the model. What I'd really like is something brilliant and minimalist (that I can't myself invent!) I know there are lots of desktop-based visualization methods that would be fascinating and could dazzle a crowd at a presentation, but I want something day-to-day useful ... and ideally something that fits on a mobile phone (or at least iPhone) screen. So that in a perfect world, if I'm out and about, I can poke this system with one piece of data and have it return the associations my brain makes -- and the ones it would make if I were jacked up on caffeine and had the whole internet in my frontal lobe somewhere. I'd like maps, charts, and pictures in there too.

I hope this outline makes some sort of sense. Like I said, it's really not as tricky or complex as it sounds. Fine tuning the metrics will take real work, as will optimizing the data structures so that the relevant queries are fast, and so that the system can work in "tinfoil-hat-private-mode" as well as "publish-whatever-you-deduce-from-my-friendfeed-and-private-chats-mode."

Incidentally, this app could also help solve the augmented reality dilemma of "how do you narrow down all the possible info about the input objects and coordinates, so that the user sees something interesting and/or actionable," so that's another angle.

Have some capital or time you want to throw in this direction? Feel free to email me -- adbreind@gmail.com ... or if you want to loot this idea and think you can build a killer app for the semantic web era? That's fine too -- send me a link when I can sign up for the beta.

Wednesday, April 23, 2008

Augmented Reality is Here, You Just Can't See It

Augmented reality is here now, and is going to get bigger fast. In fact, it's going to be a killer app for the semantic web.

Where did I come up with that?

Well, as far as not seeing it ... that's because it doesn't look like the picture you have in your head. You're thinking of stuff floating in midair like in Minority Report or at least a heads-up display with goggles or a sceen that overlays data onto video.

Those may be nice implementations. But the core utility of augmented reality is getting information on stuff you're near (where near can be physical or by mental association) and having it delivered to you in an actionable "heads-up way" even if it's not on a HUD.

In the information sense, AR is already popular. Joe calls Fred on his cell phone: "I'm trying to find a hardware store down here near 20th, can you hop on Google maps and find it?" Fred looks it up, maybe hits Street Views, done. Later he calls Fred again from a party: "That girl is here -- the one I met at your office party, used to work with you... what is her name? She's looking great. What is her whole deal again?" That's augmented reality by cellphone and human reverse TTY, call it v 0.5.

Then there's v 0.6, which is Google Local for Mobile/iPhone, Windows Live Mobile, or the like. GLM offers "My Location"-biased search results, while WLM has solid speech recognition (you can just tell it where you are or what you want). These mobile search products are great, except that they are relatively active, not passive -- you need to tell them what you're interested in, they doesn't already know. And they only knows a little bit about places and a little more about businesses.

Despite the limitations, these two methods are real examples of AR in use today, even if people don't call it that.

I assert that AR is a killer semantic web app because it's the semantic tools that let the machine filter the Internet down to what's relevant in your (metaphorical) field of view, when you're out in the real world.

Googling for answers in many real-world situations is drinking from a firehose, and you need to go all the way into the cyber world (iPhone, laptop, etc.) and invest effort to get what you want. That's not AR, it's just context switching and portable cyberspace.

AR is a system that can matrix your interests, contacts, places, and needs against all the current information germane to where you are or what you're doing ... and then pick out just the high level parts you want, with a mechanism to drill down by concept. Doing that by brute-force search, or even collaborative methods (think geotags), won't get you far enough. You need a true semantic layer to front-load the work and make this real-time.

On the other side, once you have a workable if basic semantic layer, then AR becomes a very basic incredibly useful flavor of personal, semantic search.