The NoSql squad is on the loose again, therefore it is time once again to slay that particular dragon. Andy Oram, of O'Reilly fame writes about a recent NoSql conference (prior to it). The usual suspects are listed.
The hallmark (or carbuncle) of the NoSql "databases" is that they have no structured access, no transactional support, and no data reduction. You get all those, and more, with relational (sql, mostly) databases. What these NoSql (Oram's list: Cassandra, CouchDB, HBase, HypergraphDB, Hypertable, Memcached, MongoDB, Neo4j, Riak, SimpleDB, Voldemort) datastores do is just what files did with your grandfather's COBOL code; store data in an application specific format. You can browse through them at your leisure.
What they have in common is the use of key/value pairs for data. The proponents fail to comprehend that their "data" is just what an index is in a relational database. And it's nothing new. Back in the 1960's random access was supported by the "fully inverted file" paradigm. Here's what Joe Celko has to say ("Joe Celko's Data and Databases"): "In an inverted file, every column in a table has an index on it. It is called an inverted file structure because columns become files." Nothing new here, please move on. Gad, these young-uns persist in re-inventing square wheels.
17 March 2010
13 March 2010
SQL Server does SSD
For those who do SQL Server, SQL Server Central has begun a series on SSD. This first installment is dirt basic, but there may be something useful down the line. The other side of SSC, Simple Talk, is generally first rate. I'll be following it, and recommend it.
07 March 2010
A Kindred Soul
04 March 2010
Is That the Emerald City?
Imagine my surprise. I finally found a site that marries the relational database to SSD, RethinkDB. I just found it, so I'm still a bit giddy. Goody, goody, goody. There may be intelligent life on the planet Mr. Spock.
24 February 2010
Humpty Dumpty Falls Again
Well, campers, it's happened again, STEC has crashed. As I've mentioned before, this endeavor is not primarily dedicated to stock tips, but the health of SSD vendors (and storage vendors, too) is material to our journey down the Yellow Brick Road to the Wonderful World of Oz. As I type, the pre-market value is $9.50; this past summer the share was over $40. What happened?
If you take in the message boards (Yahoo! is the one I follow), many holding shares talk of deception and corruption and such on the part of management. But, fact is, the writing was on the wall with the third quarter report, and the share tanked then as well to about $11, when they announced that EMC wouldn't be taking additional shipments in the fourth quarter. EMC said as much during their earnings. So, it was clear that STEC wasn't shipping gobs o SSD.
As I've mentioned here a number of times, the storage vendors, likely under pressure from their clients, are resisting replacing HDD racks with SSD racks on a one-for-one basis. This was predicted. Both the SSD vendors and the storage vendors have to educate the end user clients to the types of systems which will benefit from SSD, and how those applications are superior to other sorts of applications. The answer, which ought not to surprise you, dear reader, is the BCNF database. The current approach by storage vendors is to promote SSD as a caching support. That's a niche, and small at that. There won't be much future in it.
Building from the datastore, rather than the code, has not yet reached the top suite in most companies. I happened on an email the president page at HP, and sent off a missive to Hurd on just this subject, referencing their SSD promotion web page. I might even get a reply, but I won't be holding my breath. I sent along something similar to STEC. As it is used to be said, "You can't turn the Queen Mary around in a bath tub".
If you take in the message boards (Yahoo! is the one I follow), many holding shares talk of deception and corruption and such on the part of management. But, fact is, the writing was on the wall with the third quarter report, and the share tanked then as well to about $11, when they announced that EMC wouldn't be taking additional shipments in the fourth quarter. EMC said as much during their earnings. So, it was clear that STEC wasn't shipping gobs o SSD.
As I've mentioned here a number of times, the storage vendors, likely under pressure from their clients, are resisting replacing HDD racks with SSD racks on a one-for-one basis. This was predicted. Both the SSD vendors and the storage vendors have to educate the end user clients to the types of systems which will benefit from SSD, and how those applications are superior to other sorts of applications. The answer, which ought not to surprise you, dear reader, is the BCNF database. The current approach by storage vendors is to promote SSD as a caching support. That's a niche, and small at that. There won't be much future in it.
Building from the datastore, rather than the code, has not yet reached the top suite in most companies. I happened on an email the president page at HP, and sent off a missive to Hurd on just this subject, referencing their SSD promotion web page. I might even get a reply, but I won't be holding my breath. I sent along something similar to STEC. As it is used to be said, "You can't turn the Queen Mary around in a bath tub".
09 February 2010
The Three Cornered Hat
I mulled whether to write as a comment to the previous post, or to make a new one. It is a new one. I will start with an apology to Mr. Harold. It was not my intent to paint him as some sort of malfeasant human; it just so happened that his problems with an xml datastore coincidentally appeared in front of me at the same time as two other events. The primary one was the previous post discussing Phreaky Phil Factor. I thought the connection was obvious, and would be so to my bevy of regular readers. I wasn't expecting a visit from Mr. Harold; so far as I am concerned he was merely the messenger in that story. The other event is a thread on Artima, dealing the hiring/firing problem, and the ensuing discussion among Bruce Eckel, Paul English, and the group assembled, humble self included; Phreaky Phil makes an appearance there, too.
So, the point is not that some xml datastores are slow and some are fast, while some sql databases similarly. Rather, it is the comparison of paradigms among Peso's answer, the described procedural alternatives, and eXist/XQuery/xml. This has always been the problem with xml datastores, and all that redundant data, and the need to write a multi-user engine, and so forth. Those and the lingering resentment, I will acknowledge, toward Don Chamberlin for not only perverting the relational model with SQL (Dr. Codd was shut out), but also the creation of XQuery; he never really got it, having finally returned to his IMS roots. There, I feel better now.
The set based paradigm of the relational model, as implemented in an industrial strength engine, will always be faster for a generalized datastore, than looping over records in any application language. That's the line in the sand. It can be, but not necessarily be, true that a hierarchical datastore, whether IMS or xml, can have a single optimal data path for some query. In fact, that's what IMS was designed specifically to do. Such a datastore will be bad to horrid for any other query. And it will be bedeviled with all of the update anomalies well documented in the literature. There's also the issue with siloed applications (which coding against files promotes) versus shared, disciplined datastores (which mitigates against those silos).
I didn't expect Mr. Harold, or other committed xml zealots, to be suddenly overcome with the spirit of Dr. Codd, lay down their XML Spy, and join the saved. But it could happen.
So, the point is not that some xml datastores are slow and some are fast, while some sql databases similarly. Rather, it is the comparison of paradigms among Peso's answer, the described procedural alternatives, and eXist/XQuery/xml. This has always been the problem with xml datastores, and all that redundant data, and the need to write a multi-user engine, and so forth. Those and the lingering resentment, I will acknowledge, toward Don Chamberlin for not only perverting the relational model with SQL (Dr. Codd was shut out), but also the creation of XQuery; he never really got it, having finally returned to his IMS roots. There, I feel better now.
The set based paradigm of the relational model, as implemented in an industrial strength engine, will always be faster for a generalized datastore, than looping over records in any application language. That's the line in the sand. It can be, but not necessarily be, true that a hierarchical datastore, whether IMS or xml, can have a single optimal data path for some query. In fact, that's what IMS was designed specifically to do. Such a datastore will be bad to horrid for any other query. And it will be bedeviled with all of the update anomalies well documented in the literature. There's also the issue with siloed applications (which coding against files promotes) versus shared, disciplined datastores (which mitigates against those silos).
I didn't expect Mr. Harold, or other committed xml zealots, to be suddenly overcome with the spirit of Dr. Codd, lay down their XML Spy, and join the saved. But it could happen.
08 February 2010
xml News [Updated 10 March]
I have been avoiding really, really digging into the xml sarcophagous, since I'm such a warm hearted live and let live sort of guy. Well, sometimes. I know I've mentioned that my longest standing connection to blogging/sites is Cafe au Lait. The site is of Elliotte Rusty Harold, which is actually two sites, the other called Cafe con Leche dealing with things xml. While Elliotte was/is primarily a java advocate (and now writing code rather than books for a living, so far as I can tell), he has written both articles and books dealing with xml.
He and I have exchanged e-mails occasionally about xml, mostly me suggesting that he spend more time with real databases. I hadn't looked at Leche for a while, since my bookmark goes to Lait, and Leche is almost always off the bottom of the screen. But today I scrolled down, and found these two stories:
XQuery being slow
and
Xquery not doing well with errors, not that I'd advocate a java approach in a data system.
Coders still insist that they can build a robust datastore from a foundation of Lawyers' document markup language. I guess they're fans of Sarah Palin, too. And on the subject of xml suprises, Tim Bray is telling us that Oracle has decided to keep him. I've always believed that Larry wants Sun in order to take the last significant market it doesn't own: IBM mainframe clients. How having the onlie begetter of xml motivates that goal? Maybe Tim will have to larn him some relational database. Ah, cruel irony.
Update: Tim Bray announced on his blog that he's resigned from Sun, prior to "integration" in Canada, as a result he says he never will have worked for Oracle. No reason given. You can follow him at: his site.
He and I have exchanged e-mails occasionally about xml, mostly me suggesting that he spend more time with real databases. I hadn't looked at Leche for a while, since my bookmark goes to Lait, and Leche is almost always off the bottom of the screen. But today I scrolled down, and found these two stories:
XQuery being slow
and
Xquery not doing well with errors, not that I'd advocate a java approach in a data system.
Coders still insist that they can build a robust datastore from a foundation of Lawyers' document markup language. I guess they're fans of Sarah Palin, too. And on the subject of xml suprises, Tim Bray is telling us that Oracle has decided to keep him. I've always believed that Larry wants Sun in order to take the last significant market it doesn't own: IBM mainframe clients. How having the onlie begetter of xml motivates that goal? Maybe Tim will have to larn him some relational database. Ah, cruel irony.
Update: Tim Bray announced on his blog that he's resigned from Sun, prior to "integration" in Canada, as a result he says he never will have worked for Oracle. No reason given. You can follow him at: his site.
Subscribe to:
Posts (Atom)
