https://sigmodrecord.org/publications/sigmodRecord/2406/pdfs/04_Surveys_Stonebraker.pdf - обзорная статья от авторитетов индсутрии - Павло и Стоунбрейкера.
Цитата для привлечения внимания. Срывают покровы:
Вкратце:
Data Models & Query Languages
* MapReduce dead.
* Hadoop dead.
* Spark & Flink doing well.
* RocksDB ... rocks for embedded use-case.
* Для современной СУБД хорошо бы иметь API хранилища для разработки своего.
* json > xml.
* ACID это хорошо и нужно.
* wide column db are dead.
* text search engines где-то сбоку и без транзакций.
* array db интересные и хорошие но только для своих областей - научные и подобные данные натурально ложатся на модель хранения "массив".
* vector db специальный случай массива - одноразмерный. привлекают самое большое внимание разработчиков и инвесторов сегодня. Много ML use-case'ов.
* graph db - развивают совсем другую модель хранения. От этого другой API запросов. Узкоспециализированные сценарии использования.
===
System Architectures
Глава написана хорошо - даже не буду ее конспектировать - рекомендуется к прочтению полностью 🙂
Ну и в заключение нестареющие мудрости:
* Never underestimate the value of good marketing for bad products.
* Beware of DBMSs from large non-DBMS vendors.
* Do not ignore the out-of-box experience.
* Developers need to query their database directly
* The impact of AI/ML on DBMSs will be significant
Цитата для привлечения внимания. Срывают покровы:
There have been major new ideas in DBMS architectures put forward in the last two decades that reflecting changing application and hardware characteristics. These ideas range from terrific to questionable, and we
discuss them in turn.
Вкратце:
Data Models & Query Languages
* MapReduce dead.
* Hadoop dead.
* Spark & Flink doing well.
* RocksDB ... rocks for embedded use-case.
* Для современной СУБД хорошо бы иметь API хранилища для разработки своего.
* json > xml.
* ACID это хорошо и нужно.
* wide column db are dead.
* text search engines где-то сбоку и без транзакций.
* array db интересные и хорошие но только для своих областей - научные и подобные данные натурально ложатся на модель хранения "массив".
* vector db специальный случай массива - одноразмерный. привлекают самое большое внимание разработчиков и инвесторов сегодня. Много ML use-case'ов.
The key difference between vector and array DBMSs is their query patterns. The former are designed for similarity searches that find records whose vectors have the shortest distance to a given input vector in a highdimensional space.
* graph db - развивают совсем другую модель хранения. От этого другой API запросов. Узкоспециализированные сценарии использования.
SQL:2023 introduced property graph queries (SQL/PGQ) for defining and traversing graphs in a RDBMS
===
A reasonable conclusion from the above section is that non-SQL, non-relational systems are either a niche market or are fast becoming SQL/RM systems
System Architectures
Глава написана хорошо - даже не буду ее конспектировать - рекомендуется к прочтению полностью 🙂
NewSQL vendors also incorrectly anticipated that inmemory DBMS adoption would be larger in the last decade. Flash vendors drove down costs while improving storage densities, bandwidth, and latencies. Higher DRAM costs and the collapse of persistent memory(e.g., Intel Optane) means that SSDs will remain dominant for OLTP DBMSs.
The aftermath of NewSQL is a new crop of distributed, transactional SQL RDBMSs. These include TiDB [141],
CockroachDB [195], PlanetScale [60] (based on the Vitess sharding middleware [80]), and YugabyteDB [86]. The major NoSQL vendors also added transactions to their systems in the last decade despite previously strong
claims that they were unnecessary. Notable DBMSs that made the shift include MongoDB, Cassandra, and DynamoDB. This is of course due to customer requests
that transactions are in fact necessary. Google said this cogently when they discarded eventual consistency in
favor of real transactions with Spanner in 2012 [119].
At the present time, cryptocurrencies (Bitcoin) are the only use case for blockchains. In addition, there
have been attempts to build a usable DBMS on top of blockchains, notably Fluree [25], BigChainDB [12], and ResilientDB [136]. These vendors (incorrectly) promote
the blockchain as providing better security and auditability that are not possible in previous DBMSs.
Ну и в заключение нестареющие мудрости:
* Never underestimate the value of good marketing for bad products.
* Beware of DBMSs from large non-DBMS vendors.
* Do not ignore the out-of-box experience.
* Developers need to query their database directly
* The impact of AI/ML on DBMSs will be significant