Tuesday, September 22, 2009

5 Keys for Full Recovery in the Cloud

The cloud is a natural solution for disaster recovery, but careful consideration must be given before entrusting you data to a sky-high backup repository. Can you recover workloads from the cloud? How well does it scale? What's the nature of its billing system? Is its infrastructure secure? And will it offer complete protection?

While cloud computing is a familiar term, its definitions can vary greatly. So when it comes to online backup, the cloud is an important feature that can play a large role in securing and protecting during a disaster, which I like to refer to as "cloud recovery."
In order to be worthy of this cloud recovery title, a solution should have the following five features, which I have outlined below.

1. Recover Workloads in the Cloud
There is an old saying in the data protection business that the whole point of backing up is preparing to restore. Having a backup copy of your data is important, but it takes more than a pile of tapes (or an online account) to restore. You might need a replacement server, new storage, and maybe even a new data center, depending on what went wrong.
The traditional solutions to this need are to either keep spare servers in a disaster recovery data center or suffer the downtime while you order and configure new equipment. With a cloud recovery solution, you don't want just your data in the cloud -- you want the ability to actually start up applications and use them, no matter what went wrong in your environment.

2. Unlimited Scalability
If you were buying disaster recovery servers for yourself, you would have to buy one for each of your critical production servers. The whole point of recovering to the cloud is that they already have plenty of servers.
The ideal cloud recovery solution won't charge you for those servers up front but is sure to have as much capacity as you need, when you need it. Under this model, your costs are much lower than building it yourself, because you get the benefit of duplicating your environment without the cost.

3. Pay-Per-Use Billing
I love pay-as-you-go business models because they force the vendor to have a good product. Plus, this make the buying decision much easier -- just sign up for a month or two (or six), and see how it goes.
Removing the up-front price and long-term commitment shifts the risk away from the customer and onto the vendor. The vendor just has to keep the quality up to keep customers loyal.
We also know that data centers are more cost-efficient at larger scale, especially the management effort, and they require constant improvement. In your own data center, you might have some custom configurations, but in the data recovery data center, you just need racks, stacks of servers, power and cooling. You are much better off paying a monthly fee to someone who specializes.

4. Secure and Reliable Infrastructure
Lots of people like to bash cloud providers for security and reliability, but I think they hold the providers to the wrong standard. Although it is fine, in the abstract, to point out all the places where cloud providers don't achieve perfection in security and reliability, as a customer evaluating a cloud vendor, it seems better to compare them to your own capabilities.
I believe that most of the major cloud providers' infrastructures are more secure and more reliable than those of most private data centers. The point is that security and reliability are hard, but they are easier at scale. Having control over your own data center isn't enough -- you also have to spend the money to buy the necessary equipment, software , and expertise. For most companies, infrastructure is a necessary evil. Companies like Amazon (Nasdaq: AMZN) and Rackspace do infrastructure for a living, they and do it at huge scale. Sure, Amazon's outages get reported in news, but do you think you can outperform them over the next couple of years?

5. Complete Protection
Remember the "preparing to restore" line? For me, it really comes home in this idea of complete protection. If your backup product asks you what you want to protect, I am already suspicious. My vote is, "get it all." I see lots of online products offering 20GB plans, and to me, they look like an accident waiting to happen. I don't want to know which files I need to protect -- I want to click "start" and know that any time I want, I can click "recover", and there won't be any "please insert your original disk" issues.
The places people normally get bitten by this are with databases (do you have the right agent?), configuration changes (patched your server, or added a new directory of files?), and weird applications (the one that a consultant set up, and you don't really understand how it works). Complete protection means that all of these things can be protected without requiring an expert in either your own systems, or with the cloud recovery solution.

Thursday, October 16, 2008

Opera to Web developers: Come to MAMA

Opera's new Metadata Analysis and Mining Application search engine indexes data about Web site structures

Opera Software on Wednesday revealed a search engine that indexes structural information about Web pages so Web developers and standards bodies can see what technologies are being used to build Web sites and how they are being used.

The Metadata Analysis and Mining Application search engine -- "MAMA" for short -- is being tested by the company and should be released in an invitation-only beta by the end of the year, said Snorre Grimsby, vice president of quality assurance at Opera in Oslo, Norway.

MAMA grew out of tests Opera routinely does to make sure its own browser software products work well with existing Web pages that use the most commonly used Web site-creation technology, he said.

"We realized internally that we needed to be able to find lots of live sites out there that used certain technologies in certain combinations so we could test our browser on them," Grimsby said.

The resulting search engine crawls the Web, but instead of indexing the content of Web sites, as most search engines do, it discards the content and indexes the types of technologies being used on sites, such as CSS, HTML, XHTML, and the like, Grimsby said.

This information is helpful for Web developers, who can use MAMA to identify sites that are using certain kinds of technology and see how other developers have implemented it, he said.

"It's a known fact that Web developers borrow ideas from each other," Grimsby said. If developers are working with a Web application that needs, for example, a new menu system, MAMA can help them find sites that use the technology being considered to build the system to get ideas for their own implementation.

Developers also can use MAMA to see how well sites conform to current World Wide Web Consortium (W3C) specifications for commonly used Web standards, such as CSS, HTML and others. The W3C oversees the creation and maintenance of specs for many of the most prevalent Web-site development technologies.

Grimsby said that in Opera's own use of MAMA, Opera found that the average Web page has 47 discrepancies in how the site renders W3C-maintained technologies and the W3C specifications themselves.

MAMA also can be useful for the W3C and other standards bodies to help them set priorities for developing specifications. For example, if a technology is used a certain way on the majority of Web sites, or not used very much at all, the W3C "can change the spec or take something out of the spec," Grimsby said.

During an interview Wednesday, Grimsby demonstrated MAMA in real time by using it to crawl an International Data Group Web page, http://www.idg.net/idgns, to find out what technologies the site used.

According to the search engine, the site is running on version 2.2.8 of the Apache Web Server on a Windows 32-bit hardware server, has 56 hyperlinks and uses XHTML (Extensible HTML) 1.0 and CSS, he said.

In the next eight weeks Opera expects to publish a series of articles on its developer Web site about its own internal use of MAMA, noting key findings, statistics and trends the search engine discovers, he said.

By the end of the year, the company will invite key people within standards bodies to test the search engine, with a goal of releasing it publicly to developers sometime in the first or second quarter of next year, Grimsby said.