Showing posts with label Amazon. Show all posts
Showing posts with label Amazon. Show all posts

Monday, October 27, 2008

Azure -- and the Other Clouds Players -- Should Lean Forward

Since I covered Azure pretty well two weeks ago, there's not much to add except the name and the open question of which parts of the platform can be run in-house, on AMIs, or anywhere outside of MSFT data centers (via a hosting partner). And Microsoft hasn't really addressed that either (I have questions in at PDC) so the answer appears to be "not yet, stay tuned."

Now that the semi-news is out of the way, I am a little disappointed that all the cloud players haven't leaned in more, in terms of providing added-value capabilities beyond scaling. Elastic scaling is valuable, but it's a tradeoff. You are paying significantly more to be in the cloud than you would be to host equivalent compute power on own machines, or on VMs or app server instances at a consolidated host.

If you have reasonable projections about your capacity, then you're wasting money on the elasticity premium. You do get some nice operations/management capabilities ... but for apps that really need them, you still need to bring a bunch of your own, and you're taking on someone else's ops risks too.

For some businesses, these costs make sense. Here are some value-added features that would make the price persuasive for more people outside that core group:

  1. Relational and transaction capabilities. Microsoft does get the prize here, as they are the only ones offering this right now. Distributed transactions and even joins are expensive. So charge as appropriate. It's a meaningful step beyond the $/VM-CPU-cycle model that dominates now.
  2. Reverse AJAX (comet and friends). Here is a feature that is easy to describe, tricky to get right and multiplies the value of server resource elasticity. It's a perfect scenario for an established player to sell on-demand resources, and could be a differentiator in a field sorely lacking qualitative differentiation.
  3. XMPP and XMPP/BOSH (leveraging the reverse AJAX capability above). XMPP is clearly not just for IM anymore, and may evolve into the next generation transport for "web" services. Not to mention, having a big opinionated player involved may help at the next layer in the stack, namely how a payload+operation gets represented over XMPP for interop.

Those are just a couple of ideas that spring to mind -- I'm sure there are much better ones out there. To make the cloud more of a "pain killer" than a "vitamin" for more people, some new hard-to-DIY features are the way to go.

Wednesday, September 03, 2008

Amazon EC2 to Support Windows Server AMIs

Jeff Barr, Amazon Web Services Senior Evangelist, just finished giving a talk at "The AWS Start-up Event – San Francisco" in which he showed slides listing upcoming plans at AWS, including support for Windows Server.

This is an exciting development, as Windows Server / ASP.Net make for a fantastic if potentially expensive platform. Now Microsoft has to step up to the plate and come up with a pay-as-you-go, per-cycle or per-cpu-hour licensing scheme.

One of the things that makes ASP.Net interesting is that it lives in a nice middle ground between Java, which is extremely fast (in EC2, this means less expensive per transaction) and has great "enterprise" capabilities but is cumbersome to develop with, and, say, Rails, which is quite slow and has poor enterprise app cred but is very pleasant and lightweight to develop with. ASP.Net, especially with the MVC framework and forthcoming support for Python and Ruby in addition to C# and the other .Net languages, seems to combine extreme performance, easy development, and access to as much "enterprise" as you need while offering lightweight alternatives like LINQ and SSDS.

My point isn't to make a commercial for ASP.Net, but to point out that if Microsoft can get their licensing in order, they might catch up in the cloud world through fast, cheap development cycles plus faster (and hence cheaper) runtime operation on a given machine instance than some competing platforms.

Just to be fair, cloud vendor enomalism has run Windows Server on EC2 before, by virtue of the Qemu emulation software (on top of Linux). But if we're talking about maximizing efficiency, wasting cycles on another layer of emulation (EC2 instances are of course virtual to begin with) doesn't sound like the way to go.

Wednesday, August 13, 2008

Format Shifting

The dead-tree-publishing world is slowly starting to worry about online, possibly illegal, distribution of books. Not nearly enough to get smart about it though.

To me, it seems brain-dead obvious that if I buy a paper copy of a book for anywhere from $20 to $50, I ought to be able to use an electronic copy. I don't know the law on this, but I don't have pangs of conscience downloading a gray-market PDF of a book I've already bought, because the publisher wants to charge me again for an "e-book bundle" or because they make the content available in an annoying format.

But that's just the tip of the blade. In a short while, things like Kindle will take off, and the next generation -- motivated to leap in because of expensive textbooks -- will start assuming all books will be free.

Publishers should be experimenting with their models right now in attempt to adapt.

Here's one idea: the whole hardcover-softcover release cycle makes about as much sense as releasing a movie in the U.S. on Wednesday and not expecting it to be an xvid in homes in Thailand by Friday.

I like paper books, but I don't like hardcover books for anything non-classic. They're overly large, awkward to carry, just plain silly. So what happens when a publisher releases a new book I want to read in hardcover? I'm not inclined to buy it -- heck if I wanted to carry the hardcover around I could just get it out of the library.

I could download it as soon as someone puts it online, but I would want to print and bind it ... so I'd have to submit it as a print-on-demand job somewhere, hope they don't complain about the copyrighted material, and get it sent to me. I guess Kinko's is an option.

Anyone else see the problem/opportunity here?

For now, anyway, the hassle and cost of print-on-demand makes it a cost-neutral issue; I simply don't have the option getting a bound paper copy of the book for free. So given that I'm willing to -- and indeed must -- spend money to get the book, why won't the publisher sell it to me in the format I want?

Publishers ought to be taking the risks and making the friends now, before e-book readers make a big dent in the marker. Otherwise, it will be game over soon enough.

And before anyone suggests that the publishers play some critical role in blessing publications, I would point out that is just an artifact of the traditional physical distribution mechanism (paper, shelf space, etc.) ... the Internet already allows vastly more content than could be run through a press and tossed up at Borders. It's not all "good" ... but the net does a fine at creating search, moderation, and recommendation networks that allow one to find the good stuff, in a way that 3x5 "recommendation" cards tacked under a bookshelf at the local store cannot.

Wednesday, March 19, 2008

Stored Procedures and Code in the Cloud

For modest-sized Internet applications, the allure of cloud services has two elements.

First, there's simplicity of implementation and maintenance -- the hassle of real-world ops is the sort of problem startup CEOs dream of having, while startup engineers (and engineering budgets) are ill-equipped to deal with it. Second is the promise of easy scalability -- another problem the CEOs dream of, and the engineers secretly hope will become Somebody Else's Problem.

Storage in the cloud is conceptually easy. Especially with ActiveRecord patterns that ignore (at their own peril, but that's an article for another day) 35 years' worth of learnings about data integrity in the relational model. And for those who need to write things more complicated than 37signals' latest masterpiece, true structured data services in the cloud are coming.

What about application logic, though? There's raw EC2, which works on the level of provisioned VM images, and makes you design for clustering, manage your instances (while they're up), and keep your dynamic data somewhere else. Fabulous infrastructure but non-trivial to use.

Folks like heroku have value-added application-level services above EC2, which offer the elasticity with less hassle.

But what about going even higher level, and defining a unit of work, or a service module that can be deployed into a scalable container, preferable "nearby" the data it needs?

Real world example: in a recent project, I needed to be able to run Dijkstra's algorithm on big (250,000+ nodes) graphs in a persistent store. It would be great to use SimpleDB or SSDS (Astoria) for storage, but what about running the algorithm? It's not practical to extract a representation of the graph over the network, and then run Dijkstra on it just to find some interesting nodes each and every time. Changing the algorithm or using a different one? Maybe ... But what I really wanted to do was create a small module that I could ship over to the data, and run there. Even better, I'd like to be able to compute on the data "in place" in a storage facility, rather than extract. Conceptually a bit like a stored procedure.

I believe that the solution -- and an easier way to start shipping computing into a cloud facility -- is to create a module definition that one can code to, and then just upload. I think Python or Ruby would be ideal languages as they are popular, truly cross-platform, and not encumbered with a IP issues. Plus the modules would be provided as source, so that they could be scanned for, e.g., insecure or computationally intensive uses of stuff like eval. In fact, given Google's investment in Python and in interesting tooling in general, they may already be most of the way there.

I just need a better name than "stored procedures" -- that one's not getting any points for cool.

Friday, December 14, 2007

Amazon SimpleDB isn't Astoria... it Could Be, but Does it Need to Be?

A while back I wrote about Microsoft's Astoria REST-based relational data store in the cloud (or in your data center, if you want it there).

With Amazon's SimpleDB, we're a step closer to making this vision a reality. Now we're almost on track for competition (and sooner-than-later commoditization) of the new world where you don't even need MySQL to store your stuff.

Why almost? Because SimpleDB is not a full RDBMS, but is looks more like a flavor of triple store. Now, a typical (i.e., SQL-style) RDBMS can be built on top of a triple-store fairly easily. So we could will see a SQL processor, JDBC drivers, and the like, from the community pretty soon.

Another way to look at that "top layer" is to take a REST API like those used by Astoria or ActiveResource [PDF link] and simply implement that. Not as expressive as hardcore SQL, but easier, and probably enough for many applications.

What I don't see -- in the long run anyway -- is applications developing against thin wrappers specific to the Amazon triple store service itself. There's nothing fundamentally flawed in doing so ... it's just that, for a variety of reasons, data storage has evolved very slowly. The relational model is going on 40 years old, but still reigns supreme in terms of popularity, even if it has conceptual or technical flaws when put to work in today's applications.

Given the brilliant data storage alternatives that have fallen flat time and again, I doubt Amazon SimpleDB will change the way people talk about storing structured data. So SimpleDB doesn't need to be SQL but it will probably need to at least be RESTful.