Read a couple of stories this past week or so on securely erasing data but the one that caught my eye was about RunCore and their InVincible SSD. It seems they have produced a new SSD drive with internal circuitry/mechanisms for securely erasing data.
Securely erasing SSDs
Each InVincible SSD comes with a special cable with two buttons on it one for overwriting the data (intelligent destruction) and the other for destroying the NAND cells (physical destruction).
In the erase data mode (intelligent destruction), device data is overwritten on all NAND cells so that the original data is no longer readable. Presumably as this is an internal feature even over provisioned NAND cells are also overwritten. Unclear what this does to pages that are no longer programmable but perhaps they even have a way to deal with this. There was some claim that the device would be rendered to factory new condition but it seems to me that NAND endurance would still have been reduced.
In the kill NAND cells mode (physical destruction) apparently the device generates a high enough voltage internally to electronically destroy all the NAND bit cells so they are no longer readable (or writeable). Wonder if there’s any smoke that emerges when this happens.
Not sure how you insert the special cable because the device has to be powered to do any of this. It seems to me they would have been better served with an SATA diagnostic command to do the same thing, but maybe the special cable is a bit more apparent. The cable comes with two buttons one green and the other red (I would have thought yellow and red more appropriate).
But what about my other SSDs?
It’s not as useful as I first thought because what the world really needs is a device that could erase or kill NAND cells on any SSD drive. That way we could securely erase all SSDs.
I suppose the problem with a universal SSD eraser is that it would need to somehow disable wear leveling to get at over provisioned NAND cells. Also to physically destroy NAND cells would take some special circuitry. But maybe if we could come up with a standard approach across the industry such a device could be readily available.
I suppose another approach is to encrypt the data and throw away your keys but that seems to simple.
Or maybe just overwrite the data a half dozen or so times with random, repeating data patterns and then their complements. But this may not reach over-provisioned cells and with wear leveling in place all these writes could conceivably go to the same, single NAND page.
New approaches to securely erasing disk data
On another note at SNW early this year I was talking with another vendor and he said that securely erasing disk drives no longer takes multiple (3-7 depending on who you want to believe) passes of overwriting with specified data patterns (random, repeating patterns and complements of same). He said there was research done recently which had proved this but I could only find this article on [Disk] Data Sanitization.
And sometime this past week I had read another article (don’t know where) about a company shipping a device which degausses standard 3.5″ disk drives. You just insert a disk inside of it and push a button or two and your data is gone.
Why all the interest in securely erasing data?
It never really goes away. No one wants their data publicly available and securely erasing it after the fact is a simple (but lengthy) approach to deal with it.
But why isn’t everyone using data encryption? Seems like a subject for another post.
This new data center is intended to house copies of all communications intercepted the NSA. We have talked about this data center before and how it’s going to store YB of data (See my Yottabytes by 2015?! post).
One major problem with having a YB of communications intercepts is that you need to have multiple copies of it for protection in case of human or technical error.
Apparently, NSA has a secondary data center to backup its Utah facility in San Antonio. That’s one copy. We also wrote another post on protecting and indexing all this data (see my Protecting the Yottabyte Archive post)
NSA data centers
The Utah facility has enough fuel onsite to power and cool the data center for 3 days. They have a special power station to supply the 65MW of power needed. They have two side by side raised floor halls for servers, storage and switches, each with 25K square feet of floor space. That doesn’t include another 900K square feet of technical support and office space to secure and manage the data center.
In order to help collect and temporarily storage all this information, apparently the agency has been undergoing a data center building boom, renovating and expanding their data centers throughout the states. The article discusses some of other NSA information collection points/data centers, in Texas, Colorado, Georgia, Hawaii, Tennessee, and of course, Maryland.
New NSA super computers
In addition to the communication intercept storage, the article also talks about a special purpose, decrypting super computer that NSA has invented over the past decade which will also be housed in the Utah data center. The NSA seems to have created a super powerful computer that dwarfs the current best Cray XT5 super computer clusters that operate at 1.75 petaflops available today.
I suppose what with all the encrypted traffic now being generated, NSA would need some way to decrypt this information in order to understand it. I was under the impression that they were interested in the non-encrypted communications, but I guess NSA is even more interested in any encrypted traffic.
Decrypting old data
With all this data being stored, the thought is that the data now encrypted with unbreakable AES-128, -192 or -256 encryption will eventually become decypherable. At that time, foriegn government and other secret communications will all be readable.
By storing this secret communications now, they can scan this treasure trove for patterns that eventually occur and once found, such patterns will ultimately lead to decrypting the data. Now we know why they need YB of storage.
So NSA will at least know what was going on in the past. However, how soon they can move that up to do real time decryption of communications today is another question. But knowing the past, may help in understanding what’s going on today.
~~~~
So be careful what you say today even if it’s encrypted. Someone (NSA and its peers around the world) will probably be listening in and someday soon, will understand every word that’s been said.
safe 'n green by Robert S. Donovan (cc) (from flickr)
There was an minor announcement yesterday, which said something to the effect that data stored in the cloud in Europe and other locations was not immune to US Patriot act access.
This concern was mainly aired by one cloud provider but they mentioned any US company would need to provide the same access to data located anywhere.
I suppose living in the US, this sort of access should not be a concern for me but somehow this struck a chord. Does this mean that anything I store in the cloud, search on the internet, publish to social media is essentially available to any government entity that deems it important to access – yes, probably so.
The Fourth Amendment to the US constitution established the right of individuals to not be subject to “unreasonable search and seizure of property”. One could readily extend the definition of property to data. However somewhere in case law this provision has been modified to imply that such rights only apply to property that a person has a reasonable expectation of being private.
Data property rights outside your office
So where does that leave the data property rights:
Social media – seems to me that you waive any property rights to the data you submit to social media the moment you hit enter. For example, in Twitter any tweets you create are broadcast to all your followers and anybody searching on tweet text (unless you restrict your tweets) can see it. Places like Facebook, Flickr, Youtube, and other social media provide a service where updates are broadcast automatically to anyone searching on that information unless you lock it down and secure access to only a limited set of “friends”. But in the most common case, data in social media is public information (although perhaps owned by the social media company).
Cloud data – privacy rights may or may not exist in the cloud, it depends on what you store there. Lets say you start backing up your laptop/desktop to the cloud. Such data is in a format that is likely proprietary to the particular backup application you use but that doesn’t mean you have any reasonable expectation of privacy because those formats are known to the US company that created it. As such, plain text data, placed in the cloud probably has no expectation of privacy. Encrypted data is another story however.
Establishing reasonable expectations of privacy
So what can someone do today to establish “expectations of privacy”
Abandon social media. If you can’t do that, be very careful of the data you expose there.
Abandon cloud storage. If you can’t do that encrypt your data before it moves or is copied to the cloud. But you must understand who owns the encryption keys and where they reside. If the cloud provider owns the encryption keys and they can be found in the cloud, then reasonable expectation of privacy IS not present. To really secure data, encrypt the data yourself with an application not associated with the cloud service, with key phrases known only to you and stored outside the cloud only. Given all that one can assume a “reasonable expectation of privacy”.
Yes, either of these approaches are painful. Yes, they make using such facilities more complex, painful and time consuming but it’s the only way to establish a privacy rights for your data.
—-
Being an active user of Twitter and blogging, I have no reasonable expectation of privacy for this data but that doesn’t mean I relinquish the rest of my data to unrestrained access.
For some time now I have been considering the use of cloud backup but have been reluctant for my data to leave my control. Such fears, now seem to have a factual component to them. Nonetheless, cloud data can be private and secure but only if one safeguards the data before it leaves your premises.
Was invited to the SNIA tech center to witness the CDMI (Cloud Data Managament Initiative) plugfest that was going on down in Colorado Springs.
It was somewhat subdued. I always imagine racks of servers, with people crawling all over them with logic analyzers, laptops and other electronic probing equipment. But alas, software plugfests are generally just a bunch of people with laptops, ethernet/wifi connections all sitting around a big conference table.
The team was working to define an errata sheet for CDMI v1.0 to be completed prior to ISO submission for official standardization.
What’s CDMI?
CDMI is an interface standard for clients talking to cloud storage servers and provides a standardized way to access all such services. With CDMI you can create a cloud storage container, define it’s attributes, and deposit and retrieve data objects within that container. Mezeo had announced support for CDMI v1.0 a couple of weeks ago at SNW in Santa Clara.
CDMI provides for attributes to be defined at the cloud storage server, container or data object level such as: standard redundancy degree (number of mirrors, RAID protection), immediate redundancy (synchronous), infrastructure redundancy (across same storage or different storage), data dispersion (physical distance between replicas), geographical constraints (where it can be stored), retention hold (how soon it can be deleted/modified), encryption, data hashing (having the server provide a hash used to validate end-to-end data integrity), latency and throughput characteristics, sanitization level (secure erasure), RPO, and RTO.
A CDMI client is free to implement compression and/or deduplication as well as other storage efficiency characteristics on top of CDMI server characteristics. Probably something I am missing here but seems pretty complete at first glance.
SNIA has defined a reference implementations of a CDMI v1.0 server [and I think client] which can be downloaded from their CDMI website. [After filling out the “information on me” page, SNIA sent me an email with the download information but I could only recognize the CDMI server in the download information not the client (although it could have been there). The CDMI v1.0 specification is freely available as well.] The reference implementation can be used to test your own CDMI clients if you wish. They are JAVA based and apparently run on Linux systems but shouldn’t be too hard to run elsewhere. (one CDMI server at the plugfest was running on a Mac laptop).
Plugfest participants
There were a number people from both big and small organizations at SNIA’s plugfest.
Mark Carlson from Oracle was there and seemed to be leading the activity. He said I was free to attend but couldn’t say anything about what was and wasn’t working. Didn’t have the heart to tell him, I couldn’t tell what was working or not from my limited time there. But everything seemed to be working just fine.
Carlson said that SNIA’s CDMI reference implementations had been downloaded 164 times with the majority of the downloads coming from China, USA, and India in that order. But he said there were people in just about every geo looking at it. He also said this was the first annual CDMI plugfest although they had CDMI v0.8 running at other shows (i.e, SNIA SDC) before.
David Slik, from NetApp’s Vancouver Technology Center was there showing off his demo CDMI Ajax client and laptop CDMI server. He was able to use the Ajax client to access all the CDMI capabilities of the cloud data object he was presenting and displayed the binary contents of an object. Then he showed me the exact same data object (file) could be easily accessed by just typing in the proper URL into any browser, it turned out the binary was a GIF file.
The other thing that Slik showed me was a display of a cloud data object which was created via a “Cron job” referencing to a satellite image website and depositing the data directly into cloud storage, entirely at the server level. Slik said that CDMI also specifies a cloud storage to cloud storage protocol which could be used to move cloud data from one cloud storage provider to another without having to retrieve the data back to the user. Such a capability would be ideal to export user data from one cloud provider and import the data to another cloud storage provider using their high speed backbone rather than having to transmit the data to and from the user’s client.
Slik was also instrumental in the SNIA XAM interface standards for archive storage. He said that CDMI is much more light weight than XAM, as there is no requirement for a runtime library whatsoever and only depends on HTTP standards as the underlying protocol. From his viewpoint CDMI is almost XAM 2.0.
Gary Mazzaferro from AlloyCloud was talking like CDMI would eventually take over not just cloud storage management but also local data management as well. He called the CDMI as a strategic standard that could potentially be implemented in OSs, hypervisors and even embedded systems to provide a standardized interface for all data management – cloud or local storage. When I asked what happens in this future with SMI-S he said they would co-exist as independent but cooperative management schemes for local storage.
Not sure how far this goes. I asked if he envisioned a bootable CDMI driver? He said yes, a BIOS CDMI driver is something that will come once CDMI is more widely adopted.
Other people I talked with at the plugfest consider CDMI as the new web file services protocol akin to NFS as the LAN file services protocol. In comparison, they see Amazon S3 as similar to CIFS (SMB1 & SMB2) in that it’s a proprietary cloud storage protocol but will also be widely adopted and available.
There were a few people from startups at the plugfest, working on various client and server implementations. Not sure they wanted to be identified nor for me to mention what they were working on. Suffice it to say the potential for CDMI is pretty hot at the moment as is cloud storage in general.
But what about cloud data consistency?
I had to ask about how the CDMI standard deals with eventual consistency – it doesn’t. The crowd chimed in, relaxed consistency is inherent in any distributed service. You really have three characteristics Consistency, Availability and Partitionability (CAP) for any distributed service. You can elect to have any two of these, but must give up the third. Sort of like the Hiesenberg uncertainty principal applied to data.
They all said that consistency is mainly a CDMI client issue outside the purview of the standard, associated with server SLAs, replication characteristics and other data attributes. As such, CDMI does not define any specification for eventual consistency.
Although, Slik said that the standard does guarantee if you modify an object and then request a copy of it from the same location during the same internet session, that it be the one you last modified. Seems like long odds in my experience. Unclear how CDMI, with relaxed consistency can ever take the place of primary storage in the data center but maybe it’s not intended to.
—–
Nonetheless, what I saw was impressive, cloud storage from multiple vendors all being accessed from the same client, using the same protocols. And if that wasn’t simple enough for you, just use your browser.
If CDMI can become popular it certainly has the potential to be the new web file system.
MRI of my brain after surgery for Oligodendroglioma tumor by L_Family (cc) (From Flickr)
I was reading a book the other day and it suggested that sometime in the near future we will all have a personal medical record archive. Such an archive would be a formal record of every visit to a healthcare provider, with every x-ray, MRI, CatScan, doctor’s note, blood analysis, etc. that’s ever done to a person.
Such data would be our personal record of our life’s medical history usable by any future medical provider and accessible by us.
Who owns medical records?
Healthcare is unusual. For any other discipline like accounting, you provide information to the discipline expert and you get all the information you could possibly want back, to store, send to the IRS or or whatever, to do with it as you want. If you decide to pitch it, you can pretty much request a copy (at your cost) of anything for a certain number of years after the information was created.
But, in medicine, X-rays are owned and kept by the medical provider, same with MRIs, CT scans, etc. and you hardly ever get a copy. Occasionally, if the physician deems it useful for explicative reasons, you might get a grainy copy of an X-ray that shows a break or something but other than that and possible therapeutic instructions, typically nothing.
Getting Doctor’s notes is another question entirely. It’s mostly text records in some sort of database somewhere online to the medical unit. But, mainly what we get as patients, is a verbal diagnosis to take in and mull over.
Personal experience with medical records
I worked for an enlightened company a while back that had their own onsite medical practice providing all sorts of healthcare to their employees. Over time, new management decided this service was not profitable and terminated it. As they were winding down the operation, they offered to send patient medical information to any new healthcare provider or to us. Not having a new provider, I asked they send them to me.
A couple of weeks later, a big brown manilla envelope was delivered. Inside was a rather large, multy-page printout of notes taken by every medical provider I had visited throughout my tenure with this facility. What was missing from this assemblage was lab reports, x-rays and other ancillary data that was taken in conjunction with those office visits. I must say the notes were comprehensive and somewhat laden with medical terminology but they were all there to see.
Printouts were not very useful to me and probably wouldn’t be to any follow-on medical group caring for me. However the lack of x-rays, blood work, etc. might be a serious deficiency for any follow-on treatment. But, as far as I was concerned it was the first time any medical entity even offered me any information like this.
Making personal medical records useable, complete, and retrievable
To take this to the next level, and provide something useful for patients and follow-on healthcare, we need some sort of standardization of medical records across the healthcare industry. This doesn’t seem that hard, given where we are today and need not be that difficult. Standards for most medical data already exist, specifically,
DICOM or Digital Imaging and Communications in Medicine – is a standard file format used to digitally record X-Rays, MRIs, CT scans and more. Most digital medical imaging technology (except for ultrasound) out there today optionally records information in DICOM format. There just so happens to be an open source DICOM viewer that anyone can use to view these sorts of files if one is interested.
Ultrasound imaging – is typically rendered and viewed as a sort of movie and is often used for soft tissue imaging and prenatal care. I don’t know for sure but cannot find any standard like DICOM for ultrasound images. However, if they are truly movies, perhaps HD movie files would suffice for a standard ultrasound imaging file.
Audiograms, blood chemistry analysis, etc. – is provided by many technicians or labs and could all be easily represented as PDFs, scanned images, JPEG/MPEG recordings, etc. Doctors or healthcare providers often discuss salient items off these reports that are of specific interest to the patients condition. Such affiliated notes could all be in an associated text file or even a recording made of the doctor discussing the results of the analysis that somehow references the other artifact (“Blood chemistry analysis done on 2/14/2007 indicates …”).
Other doctor/healthcare provider notes – I find that everytime I visit a healthcare provider these days, they either take copious notes using WIFI connected laptops, record verbal notes to some voice recorder later transcribed into notes, or some combination of these. Any of such information could be provided in standard RTF (text files) or MPEG recordings and viewed as is.
How patients can access medical data
Most voice recordings or text notes could easily be emailed to the patient. As for DICOM images, ultrasound movies, etc., they could all be readily provided on DVDs or other removable media sent to the patient.
Another and possibly better alternative, is to have all this data uploaded to a healthcare provider’s designated URL, stored in a medical record cloud someplace, allowing patient access for viewing, downloading and/or copying. I envision something akin to a photo sharing site, upload-able by any healthcare provider but accessible for downloads by any authorized user/patient.
Medical information security
Any patient data stored in such a medical record cloud would need to be secured and possibly encrypted by a healthcare provider supplied pass code which could be used for downloading/decrypting by the patient. There are plenty of open source cryptographic algorithms which would suffice to encrypt this data (see GNU Privacy Guard for instance).
As for access passwords, possible some form of public key cryptography would suffice but it need not be that sophisticated. I prefer to use open source tools for these security mechanisms as then it would be readily available to the patient or any follow-on medical provider to access and decrypt the data.
Medical information retention period
The patient would have a certain amount of time to download these files. I lean towards months just to insure it’s done in a timely fashion but maybe it should be longer, something on the order of 7-years after a patients last visit might work. This would allow the patient sufficient time to retrieve the data and to supply it to any follow-on medical provider or stored it in their own, personal medical record archive. There are plenty of cloud storage providers I know, that would be willing to store such data at a fair, but high price, for any period of time desired.
Medical information access credentials
All the patient would need is an email and/or possible a letter that provides the accessing URL, access password and encryption passcode information for the files. Possibly such information could be provided in plaintext, appended to any bill that is cut for the visit which is sure to find its way to the patient or some financially responsible guardian/parent.
How do we get there
Bootstrapping this personal medical record archive shouldn’t be that hard. As I understand it, Electronic Medical Record (EMR) legislation in the US and elsewhere has provisions stating that any patient has a legal right to copies of any medical record that a healthcare provider has for them. If this is true, all we need do then is to institute some additional legislation that requires the healthcare provider to make those records available in a standard format, in a publicly accessible place, access controlled/encrypted via a password/passcode, downloadableby the patient and to provide the access credentials to the patient in a standard form. Once that is done, we have all the pieces needed to create the personal medical record archive I envision here.
—-
While such legislation may take some time, one thing we could all do now, at least in the US, is to request access to all medical records/information that is legally ours already. Once all the healthcare providers, start getting inundated with requests for this data, they might figure having some easy, standardized way to provide it would make sense. Then the healthcare organizations could get together and work to finalize a better solution/legislation needed to provide this in some standard way. I would think university hospitals could lead this endeavor and show us how it could be done.
I am going to a big conference next week, 2 full days out of the office. In times of yore, I would haul my trusty Macbook along and lugging it with me on both days as I move from pavilion to briefing hall, from lunch back to pavilion and from beer hall to bed.
A couple of months ago, I tried using an iPad for a different conference. I purchased an Apple Bluetooth (BT) keyboard and carried it with the iPad for most of the show. With the BT keypad, power input was just as fast as on the laptop and even faster as I didn’t need to boot anything up.
The other nice thing about the BT keyboard with the iPad is you have fine cursor controls (arrow keys) which can be used to position input pointer. I did find having to take my hand off the keyboard and touch the screen for some clicking action disconcerting and there were some iPad applications that didn’t handle the arrow keys appropriately but other than that, it worked great for power input, answering emails, and web searches.
The internal, soft iPad keyboard worked ok but wasn’t nearly as fast and didn’t support Dvorak. Also the soft keyboard in portrait mode only provides 6 lines of pages text which makes power input with feedback more difficult. In any case, I would use it to rip off quick emails, tweets, and other short stuff which worked well enough. I still took notes on paper (probably to old now to take notes on the iPad/laptop). Having the keyboard available with a moments delay, made it easy to decide to take it out to use it when I had the time or leave it in the backpack when I didn’t.
Another positive note was that the iPad took up very little desk space. Most briefing halls nowadays have these smallish retractable desk tops that can barely hold a legal pad let alone a laptop. The iPad fit these postage stamp desktops just fine.
Not sure how to quantify the weight advantage of the iPad+BT Keyboard vs. Macbook without weighing them but it is significant. Given all the junk I carry along with the laptop vs. the iPad+BT keyboard, the iPad/BT keyboard wins hands down. It’s almost like I am not carrying a computer at all.
Problems with using the iPad
There are a couple of web applications (e.g., Wordress visual editor) that seem dependent on flash to work properly, which made using the iPad to create blog posts problematic. Also, scrolling in WordPress post editor seems to be a flash application as well which made dealing with any long post edits problematic at best. Wordpress has an iPhone/iPad application which is just as good as the non-visual editor in web-based WordPress which comes in handy at these times.
Now in all honesty, I haven’t tried these in a while and these may not be flash issues as much as iPad issues. Nonetheless, I will guarantee that you will run into some websites that you use in your daily activities that use flash and won’t work. With the iPad you just will need to forego these websites and find alternatives.
In the office I am a heavy TweetDeck user. For some reason this application doesn’t work that well for the iPad. I have the latest version and all but find using Twitterific or the official Twitter App a better solution on the iPad.
I purchased the WiFi version of the iPad and iPad’s do not come with Ethernet plug-ins. Now most conference centers these days have WiFi, but it may not always work that well. Also some hotels only have WiFi in certain locations and not in the hotel rooms. All this makes having internet access somewhat sporadic. But you can always buy the 3G version if you want to and I always have my iphone for internet access in a pinch (assuming ATT has adequate conference center/hotel coverage).
I was told that the iPad power converter and connection would also charge up my 3G iPhone but this turned out not to work. Luckily, I brought along the power converter for the 3G iPhone by mistake and the cable connection between the power converter and iPad worked just fine for the iPhone. Also the cable from the power adaptor to iPad is somewhat short, so bring the extension cord in order to be able to work with the iPad while its charging.
I ended up purchasing the Apple case for the iPad. I wanted to be able to have it upright portrait or landscape while I was typing on the keyboard, have it slant upward while using the soft keypad and otherwise lie flat. The Apple iPad case does all this without problem.
Microsoft Office documents
Word documents get converted into Pages documents pretty easily but you lose all change tracking, some of the formatting, and other esoteric stuff. It’s probably ok for internal documents but I find putting together a final document using Pages still a problem. But I must say I am a novice here. Also converting Pages documents back into Word seems easy enough.
I have spent even less time with Numbers and Keynote but they seem adequate for minor stuffconvert .XLS and .PPT files to Numbers and Keynote files (but not back to .XLS and .PPT) and if I used them more probably ok for much more sophisticated work. There are other applications that seem to provide better iPhone support for Microsoft Office editing but I have yet to try them on either the iPad or iPhone. Also, beware that converting Numbers documents to Excel and Keynote to PowerPoint require Mac desktop versions of these programs.
Document availability is somewhat problematic. I met one person who emailed work documents to themselves to solve this problem. Email works ok as long as they don’t scroll out of iPad (iPad keeps the latest 200 emails max for any account which includes spam). For this purpose, I used a not-so-well-known email address and emailed my current work documents to that account. iTunes supports a way to copy files to and from the Mac or iPad which seems painless enough but the email interface worked just as well for me and I didn’t have to synch up to have the files transferred.
Beware of changing headers and footers in Pages and trying to alter them in Word once you get it back to the office. It never worked for me. I had to copy the text of the document to another fresh Word file and work the header/footers in that.
iPad security
Mac based passwords, logins, and security characteristics are a bit difficult and time-consumming to transfer to the iPad. You can manually load them in for any websites and applications you need but there is no way to transfer a whole keychain from Mac to iPad. As such, if you neglect to transfer security credentials for an important website to iPad your out of luck. Now there are some apps that profess to being able to transfer and maintain keychains on the iPhone or the iPad but I haven’t tried them yet.
Other iPad security aspects are even more problematic. The iPad can be setup to require entry of a 4 numeric character string to access it. Another setting will erase the contents of the iPad after 10 failed logins attempts. And MobileMe probably supports some way to erase an iPad that’s out of your hands (it does this for iPhones so I would think the same service would be available for the iPad but I haven’t looked into it).
But despite all that, I don’t feel the iPad is as secure as the Macbook. For one thing, I encrypt the data on the Macbook and the system password can be alphanumeric and considerably longer than 4 characters. In any case the harddrive can be removed from the Macbook but without the passkey, the data on the drive would be useless. In contrast the SSD-Flash memory on the iPad could be pulled out and analyzed without any trouble whatsoever and with proper understanding of IOS storage formatting be read in the clear.
Also the fact that its smaller and lighter it could easily be forgotten and left behind making it more lose-able. And it’s certainly more prone to being stolen because it’s smaller and lighter.
—–
At this point I will probably use the iPad for the upcoming VMworld conference just to see if it works as well the 2nd time as it did the first. It’s only two full days, what can go wrong?
Multiple cloud storage gateways either have been announced or are coming out in the next quarter or so. We have talked before about Nasuni’s file cloud storage gateway appliance, but now that more are out one can have a better appreciation of the cloud gateway space.
StorSimple
Last week I was talking with StorSimple that just introduced their cloud storage gateway which provides a iSCSI block protocol interface to cloud storage with an onsite data caching. Their appliance offers a cloud storage cache residing on disk and/or optional flash storage (SSDs) and provides iSCSI storage speeds for highly active working set data residing on the cache or cloud storage speeds for non-working set data.
Data is deduplicated to minimize storage space requirements. In addition data sent to the cloud is compressed and encrypted. Both deduplication and compression can reduce WAN bandwidth requirements considerably. Their appliance also offers snapshots and “cloud clones”. Cloud clones are complete offsite (cloud) copies of a LUN which can then be maintained in synch with the gateway LUNs by copying daily change logs and applying the logs.
StorSimple works with Microsoft’s Azure, AT&T, EMC Atmos, Iron Mountan and Amazon’s S3 cloud storage providers. A single appliance can support multiple cloud storage providers segregated on a LUN basis. Although how cross-LUN deduplication works across multiple cloud storage providers was not discussed.
Their product can be purchased as a hardware appliance with a few 100GB of NAND/Flash storage up to a 150TB of SATA storage. It also can be purchased as a virtual appliance at lower cost but also much lower performance.
Cirtas
In addition to StorSimple, I have talked with Cirtas which has yet to completely emerge from stealth but what’s apparent from their website is that the Cirtas appliance provides “storage protocols” to server systems, and can store data directly on storage subsystems or on cloud storage.
Storage protocols could mean any block storage protocol which could be FC and/or iSCSI but alternatively, it might mean file protocols I can’t be certain. Having access to independent, standalone storage arrays may mean that clients can use their own storage as a ‘cloud data cache’. Unclear how Cirtas talks to their onsite backend storage but presumably this is FC and/or iSCSI as well. And somehow some of this data is stored out on the cloud.
So from our perspective it looks somewhat similar to StorSimple with the exception that it uses external storage subsystems for its cloud data cache for Cirtas vs. internal storage for StorSimple. Few other details were publicly available as this post went out.
Panzura
Although I have not talked directly with Panzura they seem to offer a unique form of cloud storage gateway, one that is specific to some applications. For example, the Panzura SharePoint appliance actually “runs” part of the SharePoint application (according to their website) and as such, can better ascertain which data should be local versus stored in the cloud. It seems to have both access to cloud storage as well as local independent storage appliances.
In addition to a SharePoint appliance they offer a “”backup/DR” target that apparently supports NDMP, VTL, iSCSI, and NFS/CIFS protocols to store (backup) data on the cloud. In this version they show no local storage behind their appliance by which I assume that backup data is only stored in the cloud.
Finally, they offer a “file sharing” appliance used to share files across multiple sites where files reside both locally and in the cloud. It appears that cloud copies of shared files are locked/WORM like but I can’t be certain. Having not talked to Panzura before, much of their product is unclear.
In summary
We now have both a file access and at least one iSCSI block protocol cloud storage gateway, currently available, publicly announced, i.e., Nasuni and StorSimple. Cirtas, which is in the process of coming out, will support a “storage protocol” access to cloud storage and Panzura offers it all (SharePoint direct, iSCSI, CIFS, NFS, VTL & NDMP cloud storage access protocols). There are other gateways just focused on backup data, but I reserve the term cloud storage gateways for those that provide some sort of general purpose storage or file protocol access.
However, Since last weeks discussion of eventual consistency, I am becoming a bit more concerned about cloud storage gateways and their capabilities. This deserves some serious discussion at the cloud storage provider level and but most assuredly, at the gateway level. We need some sort of generic statement that says they guarantee immediate consistency for data at the gateway level even though most cloud storage providers only support “eventual consistency”. Barring that, using cloud storage for anything that is updated frequently would be considered unwise.
If anyone knows of another cloud storage gateway I would appreciate a heads up. In any case, the technology is still young yet and I would say that this isn’t the last gateway to come out but it feels like these provide coverage for just about any file or block protocol one might use to access cloud storage.
Although we have discussed securing data in the cloud before but we have not discussed IT data security in general. I count at least 6 different places one can secure IT data-at-rest today. In most cases, one has some sort of system to provide encryption/decryption services and some way to get encryption keys, generated, stored, and securely retrieved by this system. All these systems use symmetric key cryptography where the same key is used for encryption and decryption purposes. Approaches to IT data-at-rest security include data encryption performed as follows:
Drive level
Subsystem-based
Network-based
Appliance-based
HBA-based
Host-based.
Drive level encryption
For tape transports drive level encryption has been around since LTO-4 and previously with other proprietary tape formats. For disk, data encryption capabilities have been around for a long time in the consumer space and lately has been introduced into enterprise storage as well.
Encryption key management is critical to securing any drive level encryption. Key management can be supplied either externally by some sort of standalone key management software/appliance or internally from the tape library or disk subsystem controller itself.
The reasons for tape drive encryption are fairly substantial, tapes in transit can be lost or stolen. Similarly, disks can be replaced/stolen from enterprise storage subsystems and as such are subject to the same security concerns as tape volumes. As drive encryption is typically performed by special purpose hardware, it can operate with almost no overhead and thus, little impact to storage performance.
Disk subsystem-based encryption
Although there are only a few current implementations of this capability, data encryption/decryption could easily be done entirely at the subsystem level with key management available external or internal to the subsystem. Most likely this would be considered a software cryptographic solution but hardware could also be supplied to encrypt/decrypt data. With a software implementation, the impact on storage performance (especially, read back) might be considerable.
A couple of years ago, EMC, HDS and others added “secure data erasure” for disks or subsystems going out of service. However, this does nothing for operating data-at-rest security.
Network-based encryption
Both Cisco and Brocade offer data security services in the SAN or storage network facilities. Such capabilities will encrypt and decrypt data going to or from LUNs and/or tape drives. Key management can be supplied externally as well as internally to the networking equipment. Both Cisco and Brocade SAN encryption servicesare hardware encryption solutions and as such, operate at line speed with high throughput.
Appliance-based encryption
In the past, a number of companies offered appliance or standalone hardware based encryption which places the data security appliance within the data path somewhere between the host and its storage devices. Such solutions have been falling behind or recently been replaced by network based encryption solutions but still have a significant install base. Key management can be supplied internal to the appliance or externally. All appliance based encryption solutions support dedicated hardware for encryption/decryption of data.
HBA-based encryption
Last month EMC announced a new capability for their CLARiiON storage which operates in conjunction with Emulex HBAs to offer hardware HBA-based encryption for data. This solution is an interesting in that it’s almost host based, hardware solution and should have little to no impact on storage performance. Key management is supplied external to the HBA.
Host-based encryption
Host encryption has been available in the consumer and enterprise space for a number of years. Such services have seen much success with laptop data. Host based services are available from operating system vendors or special purpose applications. In the consumer space products such as PGP (recently purchased by Symantec) have been available for over a decade, similar capabilities exist in the enterprise space via special purpose “secure” file systems and other applications. Most host based cryptographic systems use software based algorithms. Although hardware host-based services are available in the mainframe, System z environment via cryptographic co-processors and the latest versions of Intel’s advanced processors with their instruction set extensions for AES encryption support.
Other data-at-rest security considerations
From a performance perspective, hardware encryption can have the least impact but it’s very expensive. In addition, drive level encryption is probably the most scaleable as the more drives you have, the more encryption throughput can be supported. Next comes the appliance or network based encryption solutions which can be scaled by purchasing more appliances or encryption blades/switches.
In contrast, software based services perform the worst but are easiest to deploy. Most consumer O/Ss support data encryption with a simple configuration change. Software solutions are the least expensive as well because there is no hardware to purchase. Software based solutions can also be scaled but only be adding more servers/subsystems.
In any event, key management cannot be overlooked for any data-at-rest security solution. Given the strength of modern day encryption algorithms, the loss of a data key is equivalent to the loss of all data encrypted with that key. So when considering key management, one should look for support of key archives, redundant key managers, key hierarchies and other advanced characteristics that make key access continuously available and disaster proof.
Data security is certainly feasible with any of these solutions. But performance, availability and ease of management must be understood before seriously considering any data-at-rest security regimin.