{"id":35216,"date":"2026-10-01T13:06:50","date_gmt":"2026-10-01T20:06:50","guid":{"rendered":"https:\/\/www.pingcap.com\/?p=35216"},"modified":"2026-10-02T13:25:09","modified_gmt":"2026-10-02T20:25:09","slug":"ai-hardware-data-architecture-at-scale","status":"publish","type":"post","link":"https:\/\/www.pingcap.com\/ko\/blog\/ai-hardware-data-architecture-at-scale\/","title":{"rendered":"The Database Is the Product: What Breaks When Memory Devices Scale"},"content":{"rendered":"<p class=\"wp-block-paragraph\"><em>Editor\u2019s note: This post originally appeared on The New Stack and is republished with permission. The original version is available\u00a0<a href=\"https:\/\/thenewstack.io\/ai-notetaker-database-architecture\/\">\uc5ec\uae30<\/a>.<\/em><\/p>\n\n\n\n<blockquote class=\"wp-block-quote is-layout-flow wp-block-quote-is-layout-flow\">\n<h2 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"Key_Takeaways\"><\/span>\ud575\uc2ec \uc694\uc57d<span class=\"ez-toc-section-end\"><\/span><\/h2>\n\n\n\n<ul class=\"wp-block-list\">\n<li>Splitting metadata (MySQL) from content (S3) with no shared transaction boundary pushes consistency problems into the application layer.<\/li>\n\n\n\n<li>At around 300 million rows, schema changes required maintenance windows, letting the database dictate the release cadence.<\/li>\n\n\n\n<li>In an AI memory product, every interaction is a database operation, so architectural debt becomes product failure.<\/li>\n\n\n\n<li>Moving its metadata layer to TiDB gave Plaud online DDL, ~10x QPS headroom, and P95 latency under 10 ms.<\/li>\n<\/ul>\n<\/blockquote>\n\n\n\n<p class=\"wp-block-paragraph\">Imagine you just finished a two-hour meeting. You were wearing a small AI work companion that promised to capture the conversation, structure it, and let you ask questions about it later. Two hours after the meeting, you open the app and ask for the transcript. You wait. The spinner turns. The whole reason you bought the device was so you\u2019d never have to remember a meeting again, and now the one thing it promised to do well \u2014 recall what was said \u2014 is the thing that\u2019s making you wait.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">This is the failure mode that fascinates me about AI hardware products, because it\u2019s not a model problem. The transcription was perfect. The summarization was good. The thing that frayed was the part nobody markets: getting the right bytes off disk and onto the screen at the moment the user asks. That\u2019s a data problem. And for a product whose entire promise is memory, a data problem is a product problem.&nbsp;<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">I want to walk through a real version of this because the team that hit it, Plaud, makes the most popular AI notetaker on the market, and the architecture that broke is the architecture almost every team in this category starts with. It feels reasonable at launch. It becomes a liability at scale. And the reasons why are worth understanding before you ship, not after.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"First_What_Was_Built_Correctly\"><\/span><strong>First, What Was Built Correctly<\/strong><span class=\"ez-toc-section-end\"><\/span><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">It would be easy to tell this story as \u201cthey made a mistake.\u201d They didn\u2019t. The original architecture was a sensible set of decisions, and I want to be precise about that before I dissect what went wrong, because the lesson is in the gap between \u201creasonable\u201d and \u201cright at scale,\u201d not in anyone\u2019s competence.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Plaud\u2019s product generates two very different kinds of data per recording. There\u2019s structured metadata: who recorded it, when, how long, what tags, what state the processing pipeline is in. And there\u2019s the unstructured payload: the audio file and its transcript, which can run tens of megabytes per session. The team did the textbook thing. They put the structured metadata in MySQL, where relational queries and transactions are cheap, and they put the large objects in S3, where storage is cheap and effectively infinite.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">If you\u2019ve built systems, you\u2019ve made this exact call. Object storage for blobs, a relational database for everything you need to query. It\u2019s in every architecture diagram. It\u2019s the default. And for a long stretch of the product\u2019s life, it worked fine.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The problem is not that the decision was wrong. The problem is that it contained a hidden assumption: that the metadata and the content could live in separate systems because they would never need to be consistent with each other in real time. For most applications, that assumption holds. For a product whose core interaction is \u201cgive me back exactly what I recorded, right now,\u201d it doesn\u2019t. And the gap between those two worlds is where everything started to fail.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"The_Transcript_Sidecar\"><\/span><strong>The Transcript Sidecar<\/strong><span class=\"ez-toc-section-end\"><\/span><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">In my article on <a href=\"https:\/\/thenewstack.io\/rag-pipeline-hybrid-search\/\">RAG retrieval<\/a>, I identified an anti-pattern I called the vector sidecar: the habit of standing up a separate vector database alongside your primary store, only to discover that the two systems can\u2019t answer a single query together. The Plaud architecture is the same shape, applied to an AI hardware product. Call it the transcript sidecar.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The transcript sidecar is what you get whenever you split structured metadata from unstructured content across two systems with no shared transaction guarantee. The metadata says a recording exists, is complete, and is ready. The content lives elsewhere, reached via a separate call, with its own latency and failure modes. Nothing ties the two together. There is no transaction boundary that spans \u201cthe row that says this transcript is ready\u201d and \u201cthe object that contains the transcript.\u201d<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">This produces three distinct compounding problems.<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li><strong>Retrieval latency is not a network problem; it\u2019s a data locality problem.<\/strong> S3 is excellent object storage. It is not a database. When you make object storage the primary retrieval path for tens-of-megabytes payloads under concurrent load, you\u2019re trading the consistency and latency guarantees of a database for the simplicity of a blob store. Under light load, the trade is invisible. Under heavy concurrent load, S3 retrieval latency and its variability become the user\u2019s experience of your product.<\/li>\n\n\n\n<li><strong>Consistency gaps open between two systems that fail independently.<\/strong> The MySQL row can indicate that a recording is ready a moment before the S3 object is durably reachable, or a replica can lag, causing the metadata a user sees and the content they fetch to disagree. With no shared transaction, the application layer inherits the job of papering over the gap, with retries, polling, and reconciliation logic that grows more elaborate every quarter.<\/li>\n\n\n\n<li><strong>The schema becomes immovable at exactly the wrong time.<\/strong> More on this below, but the short version is that the metadata store hit a scale ceiling where changing the schema required a maintenance window. For a product still evolving its feature set, that\u2019s a database governing the roadmap, rather than the other way around.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">The unifying insight is the one I keep coming back to throughout this whole series: any time you split data across systems that can fail independently, you have handed the consistency problem to the application layer, where it is not solved so much as managed, and where it compounds over time.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"When_the_Database_Governs_the_Roadmap\"><\/span><strong>When the Database Governs the Roadmap<\/strong><span class=\"ez-toc-section-end\"><\/span><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">The most interesting failure in the <a href=\"https:\/\/www.pingcap.com\/ko\/case-study\/how-plaud-eliminated-s3-latency-limitless-scale\/\">Plaud story<\/a> isn\u2019t the latency. It\u2019s the schema freeze, because it\u2019s the one that teams least expect and feel most acutely.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">With around 300 million rows, the team\u2019s MySQL setup reached a point where schema changes such as adding a column or changing an index \u2014 the routine evolution of any growing product \u2014 could no longer be performed online without risk. DDL operations on a table that large, on that architecture, meant locking, replication strain, and the real possibility of downtime. So changes had to be batched into maintenance windows.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Sit with what that means. A maintenance window for a schema change is the time the database allows the product team to ship. The roadmap now bends around the database\u2019s limitations. A feature that needs a new column will wait until the next window. The product&#8217;s cadence is set by the operational fragility of the data layer. Most teams don\u2019t notice this inversion because it happens gradually, and then one day a product manager asks why a small change is a two-week project, and the answer is \u201cthe database.\u201d<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The 300-million-row ceiling is a forcing function that teams hit later than they expect and <em>should<\/em> hit later than they do. Expect, because single-instance MySQL feels limitless right up until it doesn\u2019t. Should, because the architecture that gets you to 300 million rows is rarely the architecture that takes you past it, and the migration is far more disruptive at 300 million rows than it would have been at 30 million. The ceiling is real; it is predictable, and most teams plan for it only after they\u2019ve hit it.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"What_Changing_the_Architecture_Actually_Fixed\"><\/span><strong>What Changing the Architecture Actually Fixed<\/strong><span class=\"ez-toc-section-end\"><\/span><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Plaud\u2019s resolution was to consolidate the metadata layer into a distributed SQL database (they moved to <a href=\"https:\/\/www.tidb.io\">\ud2f0DB<\/a>), which removed the single-instance ceiling and restored online schema changes. I\u2019m going to use their numbers, but not as a product pitch. I want them as evidence for a claim about where architectural debt goes in a product like this.<\/p>\n\n\n\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><tbody><tr><td><strong>Before (dual-store-brain at scale)<\/strong><\/td><td><strong>After (consolidated metadata layer)<\/strong><\/td><\/tr><tr><td>Throughput strained under concurrent load<\/td><td>~10x QPS headroom<\/td><\/tr><tr><td>Tail latency variable under load<\/td><td>P95 under 10 ms<\/td><\/tr><tr><td>Schema changes require maintenance windows<\/td><td>Online DDL, no downtime<\/td><\/tr><tr><td>Single-instance ceiling at hundreds of millions of rows<\/td><td>Horizontal scale across multiple clusters<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<p class=\"has-text-align-center wp-block-paragraph\"><em>Figures per Plaud\u2019s reported migration results.&nbsp;<\/em><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">QPS headroom and tail latency matter, but the first result I\u2019d point to is the online DDL. Restoring the ability to change the schema without a maintenance window is what handed the roadmap back to the product team. The database stopped governing the release cadence. That\u2019s not a performance win you can put on a benchmark chart, but it\u2019s the one the product organization feels every week.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Also notice what the fix did not do: it did not attempt to store 30-megabyte audio files in the database. Large objects can still be stored in object storage. The point was never \u201cput everything in one system for its own sake.\u201d The point was that the metadata layer, the part that has to be consistent, queryable, and evolvable in real time, needed to actually behave like a database at scale, instead of becoming the fragile half of a split-brain architecture.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"Where_Architectural_Debt_Goes\"><\/span><strong>Where Architectural Debt Goes<\/strong><span class=\"ez-toc-section-end\"><\/span><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Here\u2019s the claim the Plaud numbers are evidence for. When the database is the product, architectural debt does not accumulate in some back-office system that only the platform team feels. It accumulates in the product experience, where every user feels it.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">This is the structural difference between an AI note-taker and, say, an internal analytics tool. For a product like this, every meaningful user interaction is a database operation. Recording creates rows and objects. Transcription updates state. Retrieval is a read. Editing is a write. Search is a query. There is no part of the product that isn&#8217;t, underneath, the database doing something. Architectural problems surface as product problems, one-for-one. A consistency gap becomes a transcript that is briefly missing. Slow retrieval reads as a product that is slow to remember. A frozen schema means features arrive late.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">No user will ever file a bug report that says &#8220;your metadata store and your object store lack a shared transaction boundary.&#8221; Left unaddressed, this class of debt manifests as a product that feels less dependable precisely when someone depends on it. For a product whose entire value proposition is &#8220;trust me to remember,&#8221; the debt lands on the promise itself. This is why Plaud&#8217;s decision to re-architect when it did matter. The team fixed the data layer before the failure became noticeable to users.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">This is why I say the database is the product. Not as a slogan. As a literal description of where the product\u2019s quality is determined. You can have the best transcription model in the category and still ship a product that feels unreliable, because reliability for this kind of product is a data architecture property, not a model property.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><span class=\"ez-toc-section\" id=\"The_Trap_is_Waiting_For_the_Whole_Category\"><\/span><strong>The Trap is Waiting For the Whole Category<\/strong><span class=\"ez-toc-section-end\"><\/span><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">AI hardware is having a moment. Note-takers, wearables, ambient recorders, pendants, badges \u2014 each built on the premise that the device will remember so you don\u2019t have to. The category is growing, and the products are getting better at the parts everyone talks about: the models, the form factor, the battery life.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">But underneath, almost all of them start with the same architecture Plaud started with, because it\u2019s the reasonable default. Structured metadata in a single-instance relational database. Large content in object storage. No shared transaction boundary. A schema that\u2019s easy to change at ten million rows and frozen at three hundred million. The trap is identical, and it\u2019s waiting at the same place on the growth curve for every team in the category.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The teams that will do well are the ones that recognize this early, understand that, for a memory product, the metadata layer is not back-office plumbing but the spine of the user experience, and plan their data architecture for the scale they\u2019re trying to reach rather than the scale they\u2019re at. Plaud\u2019s migration is the pattern worth studying before you hit 300 million rows, not after. The lesson is cheaper to learn from someone else\u2019s 300 million rows than from your own.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">When your database and your product are the same thing, you don\u2019t get to treat the database as someone else\u2019s problem. It is the product. Build it like one.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><em>Plaud hit the 300-million-row ceiling and re-architected its metadata layer before users ever felt it. To see how the team planned and ran that migration, <a href=\"https:\/\/www.pingcap.com\/ko\/case-study\/how-plaud-eliminated-s3-latency-limitless-scale\/\">read the full Plaud case study<\/a>.<\/em><br><\/p>","protected":false},"excerpt":{"rendered":"<p>Editor\u2019s note: This post originally appeared on The New Stack and is republished with permission. The original version is available\u00a0here. Imagine you just finished a two-hour meeting. You were wearing a small AI work companion that promised to capture the conversation, structure it, and let you ask questions about it later. Two hours after the [&hellip;]<\/p>\n","protected":false},"author":5,"featured_media":35230,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[145],"tags":[529,528,147,527,189],"class_list":["post-35216","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-thought-leadership","tag-ai-hardware","tag-data-architecture","tag-distributed-sql","tag-mysql-scaling","tag-online-ddl"],"acf":[],"featured_image_src":"https:\/\/static.pingcap.com\/files\/2026\/10\/02132142\/Copy-of-Blog-Feature.png","author_info":{"display_name":"Ed Huang","author_link":"https:\/\/www.pingcap.com\/ko\/blog\/author\/ed-huang\/"},"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v28.2 - https:\/\/yoast.com\/product\/yoast-seo-wordpress\/ -->\n<title>AI Hardware Data Architecture: What Breaks at Scale<\/title>\n<meta name=\"description\" content=\"Ed Huang on how Plaud\u2019s MySQL and S3 split hit a 300M-row ceiling, and what it teaches about AI hardware data architecture.\" \/>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/thenewstack.io\/ai-notetaker-database-architecture\/\" \/>\n<meta property=\"og:locale\" content=\"ko_KR\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"AI Hardware Data Architecture: What Breaks at Scale\" \/>\n<meta property=\"og:description\" content=\"Ed Huang on how Plaud\u2019s MySQL and S3 split hit a 300M-row ceiling, and what it teaches about AI hardware data architecture.\" \/>\n<meta property=\"og:url\" content=\"https:\/\/thenewstack.io\/ai-notetaker-database-architecture\/\" \/>\n<meta property=\"og:site_name\" content=\"TiDB\" \/>\n<meta property=\"article:publisher\" content=\"https:\/\/facebook.com\/pingcap2015\" \/>\n<meta property=\"article:published_time\" content=\"2026-10-01T20:06:50+00:00\" \/>\n<meta property=\"article:modified_time\" content=\"2026-10-02T20:25:09+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/static.pingcap.com\/files\/2026\/10\/02132142\/Copy-of-Blog-Feature.png\" \/>\n\t<meta property=\"og:image:width\" content=\"1800\" \/>\n\t<meta property=\"og:image:height\" content=\"600\" \/>\n\t<meta property=\"og:image:type\" content=\"image\/png\" \/>\n<meta name=\"author\" content=\"Ed Huang\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:creator\" content=\"@PingCAP\" \/>\n<meta name=\"twitter:site\" content=\"@PingCAP\" \/>\n<meta name=\"twitter:label1\" content=\"\uae00\uc4f4\uc774\" \/>\n\t<meta name=\"twitter:data1\" content=\"Ed Huang\" \/>\n\t<meta name=\"twitter:label2\" content=\"\uc608\uc0c1 \ub418\ub294 \ud310\ub3c5 \uc2dc\uac04\" \/>\n\t<meta name=\"twitter:data2\" content=\"11\ubd84\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\\\/\\\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\\\/\\\/thenewstack.io\\\/ai-notetaker-database-architecture\\\/#article\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/www.pingcap.com\\\/blog\\\/ai-hardware-data-architecture-at-scale\\\/\"},\"author\":{\"name\":\"Ed Huang\",\"@id\":\"https:\\\/\\\/www.pingcap.com\\\/#\\\/schema\\\/person\\\/e2787ee59fb58fadf52912ddc6e0c7ef\"},\"headline\":\"The Database Is the Product: What Breaks When Memory Devices Scale\",\"datePublished\":\"2026-10-01T20:06:50+00:00\",\"dateModified\":\"2026-10-02T20:25:09+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\\\/\\\/www.pingcap.com\\\/blog\\\/ai-hardware-data-architecture-at-scale\\\/\"},\"wordCount\":2140,\"publisher\":{\"@id\":\"https:\\\/\\\/www.pingcap.com\\\/#organization\"},\"image\":{\"@id\":\"https:\\\/\\\/thenewstack.io\\\/ai-notetaker-database-architecture\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/static.pingcap.com\\\/files\\\/2026\\\/10\\\/02132142\\\/Copy-of-Blog-Feature.png\",\"keywords\":[\"AI Hardware\",\"Data Architecture\",\"Distributed SQL\",\"MySQL Scaling\",\"Online DDL\"],\"articleSection\":[\"Thought Leadership\"],\"inLanguage\":\"ko-KR\"},{\"@type\":\"WebPage\",\"@id\":\"https:\\\/\\\/www.pingcap.com\\\/blog\\\/ai-hardware-data-architecture-at-scale\\\/\",\"url\":\"https:\\\/\\\/thenewstack.io\\\/ai-notetaker-database-architecture\\\/\",\"name\":\"AI Hardware Data Architecture: What Breaks at Scale\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/www.pingcap.com\\\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\\\/\\\/thenewstack.io\\\/ai-notetaker-database-architecture\\\/#primaryimage\"},\"image\":{\"@id\":\"https:\\\/\\\/thenewstack.io\\\/ai-notetaker-database-architecture\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/static.pingcap.com\\\/files\\\/2026\\\/10\\\/02132142\\\/Copy-of-Blog-Feature.png\",\"datePublished\":\"2026-10-01T20:06:50+00:00\",\"dateModified\":\"2026-10-02T20:25:09+00:00\",\"description\":\"Ed Huang on how Plaud\u2019s MySQL and S3 split hit a 300M-row ceiling, and what it teaches about AI hardware data architecture.\",\"breadcrumb\":{\"@id\":\"https:\\\/\\\/thenewstack.io\\\/ai-notetaker-database-architecture\\\/#breadcrumb\"},\"inLanguage\":\"ko-KR\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\\\/\\\/thenewstack.io\\\/ai-notetaker-database-architecture\\\/\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"ko-KR\",\"@id\":\"https:\\\/\\\/thenewstack.io\\\/ai-notetaker-database-architecture\\\/#primaryimage\",\"url\":\"https:\\\/\\\/static.pingcap.com\\\/files\\\/2026\\\/10\\\/02132142\\\/Copy-of-Blog-Feature.png\",\"contentUrl\":\"https:\\\/\\\/static.pingcap.com\\\/files\\\/2026\\\/10\\\/02132142\\\/Copy-of-Blog-Feature.png\",\"width\":1800,\"height\":600},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\\\/\\\/thenewstack.io\\\/ai-notetaker-database-architecture\\\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\\\/\\\/www.pingcap.com\\\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"The Database Is the Product: What Breaks When Memory Devices Scale\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\\\/\\\/www.pingcap.com\\\/#website\",\"url\":\"https:\\\/\\\/www.pingcap.com\\\/\",\"name\":\"TiDB\",\"description\":\"TiDB | SQL at Scale\",\"publisher\":{\"@id\":\"https:\\\/\\\/www.pingcap.com\\\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\\\/\\\/www.pingcap.com\\\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"ko-KR\"},{\"@type\":\"Organization\",\"@id\":\"https:\\\/\\\/www.pingcap.com\\\/#organization\",\"name\":\"PingCAP\",\"url\":\"https:\\\/\\\/www.pingcap.com\\\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"ko-KR\",\"@id\":\"https:\\\/\\\/www.pingcap.com\\\/#\\\/schema\\\/logo\\\/image\\\/\",\"url\":\"https:\\\/\\\/static.pingcap.com\\\/files\\\/2021\\\/11\\\/pingcap-logo.png\",\"contentUrl\":\"https:\\\/\\\/static.pingcap.com\\\/files\\\/2021\\\/11\\\/pingcap-logo.png\",\"width\":811,\"height\":232,\"caption\":\"PingCAP\"},\"image\":{\"@id\":\"https:\\\/\\\/www.pingcap.com\\\/#\\\/schema\\\/logo\\\/image\\\/\"},\"sameAs\":[\"https:\\\/\\\/facebook.com\\\/pingcap2015\",\"https:\\\/\\\/x.com\\\/PingCAP\",\"https:\\\/\\\/linkedin.com\\\/company\\\/pingcap\",\"https:\\\/\\\/youtube.com\\\/channel\\\/UCuq4puT32DzHKT5rU1IZpIA\"]},{\"@type\":\"Person\",\"@id\":\"https:\\\/\\\/www.pingcap.com\\\/#\\\/schema\\\/person\\\/e2787ee59fb58fadf52912ddc6e0c7ef\",\"name\":\"Ed Huang\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"ko-KR\",\"@id\":\"https:\\\/\\\/static.pingcap.com\\\/files\\\/2022\\\/10\\\/17234942\\\/avatar.jpg\",\"url\":\"https:\\\/\\\/static.pingcap.com\\\/files\\\/2022\\\/10\\\/17234942\\\/avatar.jpg\",\"contentUrl\":\"https:\\\/\\\/static.pingcap.com\\\/files\\\/2022\\\/10\\\/17234942\\\/avatar.jpg\",\"caption\":\"Ed Huang\"},\"description\":\"Co-Founder and CTO\",\"sameAs\":[\"https:\\\/\\\/www.linkedin.com\\\/in\\\/eddxhuang\\\/\"],\"url\":\"https:\\\/\\\/www.pingcap.com\\\/ko\\\/blog\\\/author\\\/ed-huang\\\/\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"AI Hardware Data Architecture: What Breaks at Scale","description":"Ed Huang on how Plaud\u2019s MySQL and S3 split hit a 300M-row ceiling, and what it teaches about AI hardware data architecture.","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/thenewstack.io\/ai-notetaker-database-architecture\/","og_locale":"ko_KR","og_type":"article","og_title":"AI Hardware Data Architecture: What Breaks at Scale","og_description":"Ed Huang on how Plaud\u2019s MySQL and S3 split hit a 300M-row ceiling, and what it teaches about AI hardware data architecture.","og_url":"https:\/\/thenewstack.io\/ai-notetaker-database-architecture\/","og_site_name":"TiDB","article_publisher":"https:\/\/facebook.com\/pingcap2015","article_published_time":"2026-10-01T20:06:50+00:00","article_modified_time":"2026-10-02T20:25:09+00:00","og_image":[{"width":1800,"height":600,"url":"https:\/\/static.pingcap.com\/files\/2026\/10\/02132142\/Copy-of-Blog-Feature.png","type":"image\/png"}],"author":"Ed Huang","twitter_card":"summary_large_image","twitter_creator":"@PingCAP","twitter_site":"@PingCAP","twitter_misc":{"\uae00\uc4f4\uc774":"Ed Huang","\uc608\uc0c1 \ub418\ub294 \ud310\ub3c5 \uc2dc\uac04":"11\ubd84"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/thenewstack.io\/ai-notetaker-database-architecture\/#article","isPartOf":{"@id":"https:\/\/www.pingcap.com\/blog\/ai-hardware-data-architecture-at-scale\/"},"author":{"name":"Ed Huang","@id":"https:\/\/www.pingcap.com\/#\/schema\/person\/e2787ee59fb58fadf52912ddc6e0c7ef"},"headline":"The Database Is the Product: What Breaks When Memory Devices Scale","datePublished":"2026-10-01T20:06:50+00:00","dateModified":"2026-10-02T20:25:09+00:00","mainEntityOfPage":{"@id":"https:\/\/www.pingcap.com\/blog\/ai-hardware-data-architecture-at-scale\/"},"wordCount":2140,"publisher":{"@id":"https:\/\/www.pingcap.com\/#organization"},"image":{"@id":"https:\/\/thenewstack.io\/ai-notetaker-database-architecture\/#primaryimage"},"thumbnailUrl":"https:\/\/static.pingcap.com\/files\/2026\/10\/02132142\/Copy-of-Blog-Feature.png","keywords":["AI Hardware","Data Architecture","Distributed SQL","MySQL Scaling","Online DDL"],"articleSection":["Thought Leadership"],"inLanguage":"ko-KR"},{"@type":"WebPage","@id":"https:\/\/www.pingcap.com\/blog\/ai-hardware-data-architecture-at-scale\/","url":"https:\/\/thenewstack.io\/ai-notetaker-database-architecture\/","name":"AI Hardware Data Architecture: What Breaks at Scale","isPartOf":{"@id":"https:\/\/www.pingcap.com\/#website"},"primaryImageOfPage":{"@id":"https:\/\/thenewstack.io\/ai-notetaker-database-architecture\/#primaryimage"},"image":{"@id":"https:\/\/thenewstack.io\/ai-notetaker-database-architecture\/#primaryimage"},"thumbnailUrl":"https:\/\/static.pingcap.com\/files\/2026\/10\/02132142\/Copy-of-Blog-Feature.png","datePublished":"2026-10-01T20:06:50+00:00","dateModified":"2026-10-02T20:25:09+00:00","description":"Ed Huang on how Plaud\u2019s MySQL and S3 split hit a 300M-row ceiling, and what it teaches about AI hardware data architecture.","breadcrumb":{"@id":"https:\/\/thenewstack.io\/ai-notetaker-database-architecture\/#breadcrumb"},"inLanguage":"ko-KR","potentialAction":[{"@type":"ReadAction","target":["https:\/\/thenewstack.io\/ai-notetaker-database-architecture\/"]}]},{"@type":"ImageObject","inLanguage":"ko-KR","@id":"https:\/\/thenewstack.io\/ai-notetaker-database-architecture\/#primaryimage","url":"https:\/\/static.pingcap.com\/files\/2026\/10\/02132142\/Copy-of-Blog-Feature.png","contentUrl":"https:\/\/static.pingcap.com\/files\/2026\/10\/02132142\/Copy-of-Blog-Feature.png","width":1800,"height":600},{"@type":"BreadcrumbList","@id":"https:\/\/thenewstack.io\/ai-notetaker-database-architecture\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/www.pingcap.com\/"},{"@type":"ListItem","position":2,"name":"The Database Is the Product: What Breaks When Memory Devices Scale"}]},{"@type":"WebSite","@id":"https:\/\/www.pingcap.com\/#website","url":"https:\/\/www.pingcap.com\/","name":"\ud2f0DB","description":"TiDB | SQL at Scale","publisher":{"@id":"https:\/\/www.pingcap.com\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/www.pingcap.com\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"ko-KR"},{"@type":"Organization","@id":"https:\/\/www.pingcap.com\/#organization","name":"PingCAP","url":"https:\/\/www.pingcap.com\/","logo":{"@type":"ImageObject","inLanguage":"ko-KR","@id":"https:\/\/www.pingcap.com\/#\/schema\/logo\/image\/","url":"https:\/\/static.pingcap.com\/files\/2021\/11\/pingcap-logo.png","contentUrl":"https:\/\/static.pingcap.com\/files\/2021\/11\/pingcap-logo.png","width":811,"height":232,"caption":"PingCAP"},"image":{"@id":"https:\/\/www.pingcap.com\/#\/schema\/logo\/image\/"},"sameAs":["https:\/\/facebook.com\/pingcap2015","https:\/\/x.com\/PingCAP","https:\/\/linkedin.com\/company\/pingcap","https:\/\/youtube.com\/channel\/UCuq4puT32DzHKT5rU1IZpIA"]},{"@type":"Person","@id":"https:\/\/www.pingcap.com\/#\/schema\/person\/e2787ee59fb58fadf52912ddc6e0c7ef","name":"Ed Huang","image":{"@type":"ImageObject","inLanguage":"ko-KR","@id":"https:\/\/static.pingcap.com\/files\/2022\/10\/17234942\/avatar.jpg","url":"https:\/\/static.pingcap.com\/files\/2022\/10\/17234942\/avatar.jpg","contentUrl":"https:\/\/static.pingcap.com\/files\/2022\/10\/17234942\/avatar.jpg","caption":"Ed Huang"},"description":"Co-Founder and CTO","sameAs":["https:\/\/www.linkedin.com\/in\/eddxhuang\/"],"url":"https:\/\/www.pingcap.com\/ko\/blog\/author\/ed-huang\/"}]}},"grav_blocks":false,"card_markup":"<a class=\"card-resource bg-white\" href=\"https:\/\/www.pingcap.com\/ko\/blog\/ai-hardware-data-architecture-at-scale\/\"><div class=\"card-resource__image-container\"><img class=\"card-resource__image\" alt=\"Copy of Blog - Feature\" src=\"https:\/\/static.pingcap.com\/files\/2026\/10\/02132142\/Copy-of-Blog-Feature.png\" loading=\"lazy\" width=1800 height=600 \/><\/div><div class=\"card-resource__content-container\"><div class=\"card-resource__content-head\"><div class=\"card-resource__category\">Thought Leadership<\/div><\/div><h5 class=\"card-resource__title\">The Database Is the Product: What Breaks When Memory Devices Scale<\/h5><\/div><\/a>","_links":{"self":[{"href":"https:\/\/www.pingcap.com\/ko\/wp-json\/wp\/v2\/posts\/35216","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.pingcap.com\/ko\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.pingcap.com\/ko\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.pingcap.com\/ko\/wp-json\/wp\/v2\/users\/5"}],"replies":[{"embeddable":true,"href":"https:\/\/www.pingcap.com\/ko\/wp-json\/wp\/v2\/comments?post=35216"}],"version-history":[{"count":9,"href":"https:\/\/www.pingcap.com\/ko\/wp-json\/wp\/v2\/posts\/35216\/revisions"}],"predecessor-version":[{"id":35235,"href":"https:\/\/www.pingcap.com\/ko\/wp-json\/wp\/v2\/posts\/35216\/revisions\/35235"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.pingcap.com\/ko\/wp-json\/wp\/v2\/media\/35230"}],"wp:attachment":[{"href":"https:\/\/www.pingcap.com\/ko\/wp-json\/wp\/v2\/media?parent=35216"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.pingcap.com\/ko\/wp-json\/wp\/v2\/categories?post=35216"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.pingcap.com\/ko\/wp-json\/wp\/v2\/tags?post=35216"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}