<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>Data Archives - ShiftMag</title>
	<atom:link href="https://shiftmag.dev/category/data/feed/" rel="self" type="application/rss+xml" />
	<link>https://shiftmag.dev/category/data/</link>
	<description>Insightful engineering content &#38; community</description>
	<lastBuildDate>Fri, 31 Oct 2025 09:47:20 +0000</lastBuildDate>
	<language>en-GB</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.0.3</generator>

<image>
	<url>https://shiftmag.dev/wp-content/uploads/2024/08/cropped-ShiftMag-favicon-32x32.png</url>
	<title>Data Archives - ShiftMag</title>
	<link>https://shiftmag.dev/category/data/</link>
	<width>32</width>
	<height>32</height>
</image> 
	<item>
		<title>The Day Our Data Center Went Ghost</title>
		<link>https://shiftmag.dev/the-day-our-data-center-went-ghost-6756/</link>
		
		<dc:creator><![CDATA[Leo Tausanovic]]></dc:creator>
		<pubDate>Fri, 31 Oct 2025 09:47:20 +0000</pubDate>
				<category><![CDATA[Data]]></category>
		<category><![CDATA[database]]></category>
		<category><![CDATA[Halloween 2025]]></category>
		<guid isPermaLink="false">https://shiftmag.dev/?p=6756</guid>

					<description><![CDATA[<p>It’s Halloween. Want to read a horror story? This one’s set in a data center.</p>
<p>The post <a href="https://shiftmag.dev/the-day-our-data-center-went-ghost-6756/">The Day Our Data Center Went Ghost</a> appeared first on <a href="https://shiftmag.dev">ShiftMag</a>.</p>
]]></description>
										<content:encoded><![CDATA[<figure class="wp-block-post-featured-image"><img fetchpriority="high" decoding="async" width="1200" height="630" src="https://shiftmag.dev/wp-content/uploads/2025/10/houted-data-center.png?x32039" class="attachment-post-thumbnail size-post-thumbnail wp-post-image" alt="" style="object-fit:cover;" srcset="https://shiftmag.dev/wp-content/uploads/2025/10/houted-data-center.png 1200w, https://shiftmag.dev/wp-content/uploads/2025/10/houted-data-center-300x158.png 300w, https://shiftmag.dev/wp-content/uploads/2025/10/houted-data-center-1024x538.png 1024w, https://shiftmag.dev/wp-content/uploads/2025/10/houted-data-center-768x403.png 768w" sizes="(max-width: 1200px) 100vw, 1200px" /></figure>


<p class="wp-block-paragraph">It was a <strong>great vacation</strong>, packed with long walks, beach time, fun, and rest.</p>



<p class="wp-block-paragraph">I came back to work rejuvenated, knowing I’d need a few days to get back into the working rhythm. Most of the day was spent catching up with colleagues and the latest news.</p>



<p class="wp-block-paragraph">One was about a <strong>brand-new data center</strong>, ready to go live in just a few days.</p>



<h2 class="wp-block-heading">&#8220;Production-ready&#8221; turned into &#8220;data-empty&#8221;</h2>



<p class="wp-block-paragraph">&#8220;It turns out delivering new data centers is <strong>becoming almost routine</strong>,&#8221; I thought to myself, feeling as light as a breeze.</p>



<p class="wp-block-paragraph">After catching up with my colleagues, I was greeted by a mountain of unread emails. Honestly, I can’t tackle that for hours without feeling my will to live slowly evaporate. So, I decided to shift my focus and check out the new data center, just to see how everything was set up…</p>



<p class="wp-block-paragraph">My fresh bronze tan didn’t help &#8211; I went pale as a ghost: <strong>one of the key databases in our shiny new “production-ready” data center was completely empty</strong>.</p>



<p class="wp-block-paragraph">No data. Nothing. Niente. Nada.</p>



<p class="wp-block-paragraph">The problem? We needed a full month of historical data available FROM DAY ONE. And all we got were… good vibes and positive energy.</p>



<h2 class="wp-block-heading"><span id="two-days-no-data-one-hero">Two days, no data, one hero</span></h2>



<p class="wp-block-paragraph">The clock was ticking. We had maybe two days (probably less) until go-live. Yes, it was an emergency, and no, there wasn’t a second to waste.</p>



<p class="wp-block-paragraph">We turned to our <strong>trusty old Replay app</strong>, hoping it could save the day.</p>



<p class="wp-block-paragraph">And once again &#8211; it delivered! </p>



<p class="wp-block-paragraph">The shining star of our toolkit, Replay came through, <strong>regenerating over a month’s worth of data</strong> before the deadline.</p>



<p class="wp-block-paragraph">In the end, the data was restored, the deadline met, and our sanity… mostly intact. But one thing was clear: returning from vacation had never felt more like a horror story. And you? <strong>Keep your tools close and your backups closer</strong>.</p>
<p>The post <a href="https://shiftmag.dev/the-day-our-data-center-went-ghost-6756/">The Day Our Data Center Went Ghost</a> appeared first on <a href="https://shiftmag.dev">ShiftMag</a>.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Devs and DBAs can’t find peace, but could they call a truce?</title>
		<link>https://shiftmag.dev/devs-and-dbas-relationship-4930/</link>
		
		<dc:creator><![CDATA[Toni Babic]]></dc:creator>
		<pubDate>Fri, 18 Apr 2025 13:12:40 +0000</pubDate>
				<category><![CDATA[Data]]></category>
		<category><![CDATA[database]]></category>
		<category><![CDATA[Database admins]]></category>
		<category><![CDATA[DBA]]></category>
		<category><![CDATA[Developer Productivity]]></category>
		<category><![CDATA[development]]></category>
		<guid isPermaLink="false">https://shiftmag.dev/?p=4930</guid>

					<description><![CDATA[<p>Are DBAs the guardians of order or just here to give devs a hard time? Or maybe devs are a little too used to getting their way?</p>
<p>The post <a href="https://shiftmag.dev/devs-and-dbas-relationship-4930/">Devs and DBAs can’t find peace, but could they call a truce?</a> appeared first on <a href="https://shiftmag.dev">ShiftMag</a>.</p>
]]></description>
										<content:encoded><![CDATA[<figure class="wp-block-post-featured-image"><img decoding="async" width="1200" height="630" src="https://shiftmag.dev/wp-content/uploads/2025/02/devs-vs-dbs.png?x32039" class="attachment-post-thumbnail size-post-thumbnail wp-post-image" alt="" style="object-fit:cover;" srcset="https://shiftmag.dev/wp-content/uploads/2025/02/devs-vs-dbs.png 1200w, https://shiftmag.dev/wp-content/uploads/2025/02/devs-vs-dbs-300x158.png 300w, https://shiftmag.dev/wp-content/uploads/2025/02/devs-vs-dbs-1024x538.png 1024w, https://shiftmag.dev/wp-content/uploads/2025/02/devs-vs-dbs-768x403.png 768w" sizes="(max-width: 1200px) 100vw, 1200px" /></figure>


<p class="wp-block-paragraph">Picture this scenario: You&#8217;ve just joined a new tech company, excited to start onboarding and dive into your first projects. <strong>But soon, it hits you &#8211; the tension</strong>.</p>



<p class="wp-block-paragraph">Developers gripe about sluggish and obstructive DBAs, while Database Administrators grumble and murmur about the constant stream of poorly optimized queries, chaotic code, and last-minute demands from developers.</p>



<p class="wp-block-paragraph">Did you stumble upon a toxic workplace? Not exactly. <strong>It&#8217;s</strong> <strong>the age-old clash that&#8217;s been brewing for decades between two &#8216;tribes&#8217;</strong>: those who want to build quickly and those who safeguard stability.</p>



<p class="wp-block-paragraph">So, do DBAs play the role of protectors, or are devs just a tad bit spoiled?</p>



<h2 class="wp-block-heading"><span id="how-did-we-get-here">How did we get here?</span></h2>



<p class="wp-block-paragraph">The friction comes from a classic disconnect: <strong>speed vs. stability</strong>.</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">Developers are eager to ship features as fast as they can (preferably yesterday), often under pressure from clients, management, or some other deadline-driven force, while DBAs are focused on ensuring the database doesn’t crash tomorrow.</p>
</blockquote>



<p class="wp-block-paragraph">When these priorities clash, chaos is inevitable &#8211; and I know this firsthand because I’ve been on both sides.</p>



<h2 class="wp-block-heading"><span id="let%e2%80%99s-take-a-closer-look-at-it-from-a-developer%e2%80%99s-angle">Let’s take a closer look at it from a developer’s angle</span></h2>



<ul class="wp-block-list">
<li><strong>‘</strong><strong>We need to speed up our delivery!’</strong></li>
</ul>



<p class="wp-block-paragraph">The main argument is that developers often work under tight deadlines. <strong>To them, a database is just a tool</strong> &#8211; the simpler, the better. Why involve DBAs early if it slows down prototyping?</p>



<p class="wp-block-paragraph">There’s also a prevailing ‘good enough for now’ mentality, where getting a working prototype is the priority.</p>



<ul class="wp-block-list">
<li><strong>‘</strong><strong>Setting up new databases takes ages</strong><strong>.</strong><strong>’</strong></li>
</ul>



<p class="wp-block-paragraph">Waiting for approvals, scheme checks, or infrastructure setups feels like watching paint dry.</p>



<ul class="wp-block-list">
<li><strong>‘</strong><strong>Getting DBAs involved can feel like pulling teeth</strong><strong>.’</strong></li>
</ul>



<p class="wp-block-paragraph">When DBAs push back on quick fixes or demand documentation, developers see red tape rather than guardianship.</p>



<h2 class="wp-block-heading"><span id="now-let%e2%80%99s-dive-into-the-dba%e2%80%99s-side-of-the-story">Now, let’s dive into the DBA’s side of the story</span></h2>



<ul class="wp-block-list">
<li><strong>‘We’re called in too late!’</strong></li>
</ul>



<p class="wp-block-paragraph">DBAs often inherit <strong>poorly designed schemas or performance nightmares</strong>. Fixing a burning database at 2 a.m. is never fun, especially when you don’t know anything about it or where to begin.</p>



<ul class="wp-block-list">
<li><strong>‘Devs don’t listen.’</strong></li>
</ul>



<p class="wp-block-paragraph">Recommendations about indexing, query optimization, or security often fall on deaf ears &#8211; <strong>until an outage happens</strong>. No matter how many times you share documentation, hold knowledge-sharing sessions, or discuss the topic, it sometimes feels like no one listens.</p>



<ul class="wp-block-list">
<li><strong>‘We’re the unsung janitors.’</strong></li>
</ul>



<p class="wp-block-paragraph">While devs chase innovation, DBAs clean up the mess. Their priorities &#8211; backups, scalability, and compliance &#8211; may not be glamorous, but they’re non-negotiable.</p>



<p class="wp-block-paragraph">If a company wants to make it in the worldwide market, it needs to be compliant with numerous regulations. And that is just the beginning &#8211; <strong>data needs to be secured</strong>, and data leaks and privacy concerns aren’t fun for anyone. We need backups and replication set up for disaster recovery, satisfy numerous architecture decision records, and all that complicates the process. It will never be as simple and fast as just deploying a DB locally in Docker.</p>



<h2 class="wp-block-heading"><span id="so-what-do-devs-and-dbas-really-want">So, what do devs and DBAs really want?</span></h2>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">DBAs just want <strong>preventative care</strong> &#8211; they’re like doctors urging you to eat your veggies, not just treating heart attacks.<br></p>



<p class="wp-block-paragraph">On the other hand, devs need <strong>speed without fallout</strong>. They want to innovate without waking up to a dumpster fire.</p>
</blockquote>



<h2 class="wp-block-heading"><span id="but-do-devs-really-want-to-own-databases">But do devs really want to own databases?</span></h2>



<p class="wp-block-paragraph">Sure, devs can deploy their own database, but maintaining performance, security, and costs? <strong>That’s a full-time job</strong> &#8211; one many devs aren’t signing up for.</p>



<p class="wp-block-paragraph">It may sound great at first, but once you deploy it, the database works fine &#8211; until it doesn’t. It’s just a matter of time before the first issues arise. And surely, devs would rather spend their time programming than maintaining and troubleshooting a database.</p>



<h2 class="wp-block-heading">Think it can&#8217;t get worse? Well, say hello to AI</h2>



<p class="wp-block-paragraph">Enter AI and vector databases &#8211; tools promising lightning-fast analytics.</p>



<p class="wp-block-paragraph">The number of new vector databases is skyrocketing, and devs want them all, each with a specific feature they need. But for DBAs, this is a nightmare. <strong>Integrating a new database is a slow and painful process</strong>, as all checks for security, compliance, and requirements need to be tested and approved. New database types (e.g., vector DBs for ML) often lack mature tooling or expertise. DBAs scramble to secure and scale them, while devs resent the learning curve.</p>



<h2 class="wp-block-heading"><span id="so-who-is-right">So, who is right?</span></h2>



<p class="wp-block-paragraph">They both are, and at the same time, they both aren’t &#8211; <strong>much like the relationship between devs and clients</strong>.</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">Devs often complain about clients&#8217; broad demands, tight deadlines, and constant changes. Sound familiar? The Dev vs DBA battle is similar: devs forget that, as clients of DBAs, they’re the ones pushing for deadlines and changes, just like clients do to them.</p>
</blockquote>



<p class="wp-block-paragraph">So, devs must make compromises with clients because the company&#8217;s revenue depends on it. However, I think DBAs tend to be &#8216;harder&#8217; on their clients since they don’t face the same pressure of potentially losing business.</p>



<h2 class="wp-block-heading"><span id="is-there-hope-for-peace">Is there hope for peace?</span></h2>



<p class="wp-block-paragraph">There isn&#8217;t &#8211; at least not a simple one.</p>



<p class="wp-block-paragraph">But we can all work together to reduce tension and aim for some kind of truce. We need to be more understanding.</p>



<p class="wp-block-paragraph">Devs should recognize that <strong>DBAs aren’t just complicating things for no reason</strong> and should involve them in the service architecture process. On the other hand, <strong>DBAs should view devs as clients</strong>, understand the need for compromise, and simplify the process of providing databases by using automation or internal tools whenever possible.</p>
<p>The post <a href="https://shiftmag.dev/devs-and-dbas-relationship-4930/">Devs and DBAs can’t find peace, but could they call a truce?</a> appeared first on <a href="https://shiftmag.dev">ShiftMag</a>.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>From file frustration to streamlined data exchange: Rethinking the approach</title>
		<link>https://shiftmag.dev/from-file-frustration-to-streamlined-data-exchange-999/</link>
		
		<dc:creator><![CDATA[Alen Kosanovic]]></dc:creator>
		<pubDate>Tue, 25 Jul 2023 17:29:06 +0000</pubDate>
				<category><![CDATA[Backend]]></category>
		<category><![CDATA[Data]]></category>
		<category><![CDATA[anti-pattern]]></category>
		<category><![CDATA[CSV]]></category>
		<category><![CDATA[data exchange]]></category>
		<guid isPermaLink="false">https://shiftmag.dev/?p=999</guid>

					<description><![CDATA[<p>Each time some data is exchanged using files, a red flag should pop up, and you should ask yourself – what is a better way to do this? There probably is one.</p>
<p>The post <a href="https://shiftmag.dev/from-file-frustration-to-streamlined-data-exchange-999/">From file frustration to streamlined data exchange: Rethinking the approach</a> appeared first on <a href="https://shiftmag.dev">ShiftMag</a>.</p>
]]></description>
										<content:encoded><![CDATA[<figure class="wp-block-post-featured-image"><img decoding="async" width="1200" height="630" src="https://shiftmag.dev/wp-content/uploads/2023/07/data_exchange.png?x32039" class="attachment-post-thumbnail size-post-thumbnail wp-post-image" alt="" style="object-fit:cover;" srcset="https://shiftmag.dev/wp-content/uploads/2023/07/data_exchange.png 1200w, https://shiftmag.dev/wp-content/uploads/2023/07/data_exchange-300x158.png 300w, https://shiftmag.dev/wp-content/uploads/2023/07/data_exchange-1024x538.png 1024w, https://shiftmag.dev/wp-content/uploads/2023/07/data_exchange-768x403.png 768w" sizes="(max-width: 1200px) 100vw, 1200px" /></figure>


<p class="wp-block-paragraph">Occasionally, everyone needs to handle data stored in files. CSV files seem like the way to go and are often the first choice &#8211; but I&#8217;d rather call them an anti-pattern.</p>



<p class="wp-block-paragraph">Let&#8217;s say a customer requests an Excel report containing specific data. If this is a one-time thing, we can conveniently send it via email. In most cases, however, the data needs to be sent regularly.&nbsp;Not all clients use sophisticated software like Kafka; even when they do, we don&#8217;t always have access to it. So we need a solution that is both universally available and easily understood. SFTP and CSV files seemingly fit the bill.</p>



<p class="wp-block-paragraph">This solution seems optimal for sending data to an external client. The problem arises when an internal application needs to integrate with ours, requires some data, and we decide to solve it by exporting CSVs. At first glance, this seems a low-hanging fruit.</p>



<p class="wp-block-paragraph">Everyone can download a file.</p>



<p class="wp-block-paragraph">Everyone can read a CSV.</p>



<p class="wp-block-paragraph">It&#8217;s already developed.</p>



<p class="wp-block-paragraph"><strong>While it may appear&nbsp;<em>that</em>&nbsp;simple, ultimately, it proves to be more costly than using a dedicated component for data exchange. </strong>I learned my lesson the hard way. I&#8217;ve worked on both applications that sent and received data via files. Here are some of the issues I discovered along the way.</p>



<h2 class="wp-block-heading"><span id="data-storage">&nbsp;<strong>Data storage</strong></span></h2>



<p class="wp-block-paragraph">The first thing we need is a place to store the data. Since our data is stored in files, it must be placed within a file system. The decision we face is whether to store the files on the&nbsp;<strong>producer</strong>&nbsp;machine,&nbsp;<strong>consumer</strong>&nbsp;machine, or an&nbsp;<strong>intermediary</strong>&nbsp;machine. Each option has its advantages and disadvantages.</p>



<p class="wp-block-paragraph"><strong>Using the producer</strong>&nbsp;machine results in the consumer bearing the entire networking load. The consumer will need to periodically retrieve the data through a&nbsp;<strong>PULL</strong>&nbsp;operation. In the event of a network link failure, the producer will remain unaffected.&nbsp;</p>



<p class="wp-block-paragraph">However, this solution has a drawback &#8211; when consumers are redundant, and there are multiple instances, each instance may attempt to download the same file. Consequently, the same data will be processed multiple times. To address this,&nbsp;<strong>each instance should reserve the file before initiating the download</strong>. This can be accomplished by appending a suffix to the file name or moving the file to another folder. Once the file is reserved, the consumer application can proceed with the file download.</p>



<p class="wp-block-paragraph"><strong>Using the consumer</strong>&nbsp;machine puts all the networking burden on the producer. The producer needs to&nbsp;<strong>PUSH&nbsp;</strong>the data to the consumer machine once the file is created. The consumer machine can watch the local file system for changes and react when a new file is created. As a result, the push-based approach can yield a faster response by the consumer and does not put any networking burden on the consumer. Since the file is sent directly to the consumer instance, this approach does not need&nbsp;<em>file reservation.</em>&nbsp;</p>



<p class="wp-block-paragraph">Pushing files to consumers also has a considerable downside &#8211; the producer must balance these files between the consumer instances. Implementing effective client-side balancing is a tough nut to crack. The producer would need to adapt its balancing strategy when a consumer instance gets offline or online and is added to or removed from the group.</p>



<p class="wp-block-paragraph"><strong>Using an intermediary machine</strong>&nbsp;combines elements of both previously mentioned approaches. It eliminates the need for client-side balancing, but the consumer still needs to reserve the file before downloading it. The data storage is independent of both the consumer and producer, although both parties are still required to transfer files over the network.</p>



<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="1024" height="538" src="https://shiftmag.dev/wp-content/uploads/2023/07/data-exchange1-1024x538.png?x32039" alt="" class="wp-image-1053" srcset="https://shiftmag.dev/wp-content/uploads/2023/07/data-exchange1-1024x538.png 1024w, https://shiftmag.dev/wp-content/uploads/2023/07/data-exchange1-300x158.png 300w, https://shiftmag.dev/wp-content/uploads/2023/07/data-exchange1-768x403.png 768w, https://shiftmag.dev/wp-content/uploads/2023/07/data-exchange1.png 1200w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /></figure>



<h2 class="wp-block-heading"><span id="data-retention"><strong>Data retention</strong></span></h2>



<p class="wp-block-paragraph">Regardless of the chosen method for storing data, it is crucial to ensure that the disk is not filled with files. Failing to do so will render the machines (and the services running on them) unusable. The primary line of defense is to&nbsp;<strong>store files on a separate section of the disk</strong>.&nbsp;</p>



<p class="wp-block-paragraph">Even if the partition gets filled with files, it will not impact other services. This advice applies even when files are stored on an intermediary machine, as some operating system functionalities may not work properly if the disk space is occupied.</p>



<p class="wp-block-paragraph">Accumulating files can happen quite easily. For instance, the consumer application may be down (or not deployed) while the producer application continues to generate files. It can also occur if the producer generates a high volume of data the consumer cannot keep up with.&nbsp;</p>



<p class="wp-block-paragraph">In such cases, it is necessary to prevent the files from <strong>overflowing the disk.</strong> If the data represents events and skipping a few events in the event of a failure is acceptable, a retention policy should be implemented.&nbsp;</p>



<p class="wp-block-paragraph">For example, a script that periodically deletes the oldest files can be developed. This script adds <strong>another component to the architecture that needs to be created, maintained, and monitored</strong>. If the script malfunctions or contains a bug, the data retention may not work as intended.</p>



<h2 class="wp-block-heading"><span id="data-processing"><strong>Data processing</strong></span></h2>



<p class="wp-block-paragraph">A file is removed from the disk when a consumer has read and processed the whole file content. The &#8216;whole file content&#8217; is emphasized since a file represents a batch of data. Errors or failures can happen at any point between parsing and processing the data. If you delete the file without processing all the data contained within it, the data will be lost. On the other hand, if you do not delete the file but have already processed some of the data, you will end up processing the same data twice. This poses a data integrity concern that needs to be addressed. Both discarding data and processing data twice are undesirable and potentially unacceptable. That is why&nbsp;<strong>it is necessary to process a file in its entirety before deleting it</strong>.</p>



<p class="wp-block-paragraph">This can be challenging, especially if the application involves multiple processing steps, including several database writes. Errors occurring in any of these phases can result in the entire batch being reversed. Additionally, processing a batch of data within a single transaction restricts the data processing to only one thread. File reading and data processing should be performed in separate threads to optimize performance.</p>



<p class="wp-block-paragraph">One possible solution is to relax the data integrity requirement by&nbsp;<strong>implementing a queue where all the data is stored after parsing</strong>. This approach considers a file processed once all its records are inserted into the queue. Records can then be dequeued one by one, preventing the entire batch from being reverted. The issue with this approach is that the data in the queue will be lost in the event of an abrupt application failure. Since abrupt application failures are expected to be rare, this trade-off may be acceptable. If preserving data is not acceptable, a persistent queue for the data can be introduced (but we opted out of using new components in favor of using&nbsp;<em>simple CSV reading</em>).</p>



<p class="wp-block-paragraph">It is also important to&nbsp;<strong>set a reasonable limit to the number of records in a file</strong>. A large file might present a huge load of data at once. So huge that it might crash the data if it does not have the memory to read the whole file.</p>



<h2 class="wp-block-heading"><span id="data-sharing"><strong>Data Sharing</strong></span></h2>



<p class="wp-block-paragraph">Once everything is set up and all the monitoring points are implemented (excluded from this article), we can enjoy our manually implemented file-based data exchange solution… until a new consumer is needed for the same data.&nbsp;</p>



<p class="wp-block-paragraph">Fortunately, we already have this functionality (file download+reservation, data retention, data queuing, monitoring, etc.) supported, and we can easily re-use it in any new consumer. Hopefully, the code has been written well enough to be copied and pasted into a library. Otherwise, there will be significant code duplication. Let&#8217;s not forget to create<strong> documentation for the module </strong>so that the solution can be integrated effortlessly.</p>



<p class="wp-block-paragraph">Once we move this code to a module, our file-based data exchange solution will become a reusable component that any application can utilize (provided it is written in the same programming language). The only thing left is instructing the producer to export the files to a new location. There might be some duplicated data between the two locations, but let&#8217;s hope we will not add many of these consumers to the system.</p>



<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="1024" height="538" src="https://shiftmag.dev/wp-content/uploads/2023/07/data-exchange-files-1024x538.png?x32039" alt="" class="wp-image-1052" srcset="https://shiftmag.dev/wp-content/uploads/2023/07/data-exchange-files-1024x538.png 1024w, https://shiftmag.dev/wp-content/uploads/2023/07/data-exchange-files-300x158.png 300w, https://shiftmag.dev/wp-content/uploads/2023/07/data-exchange-files-768x403.png 768w, https://shiftmag.dev/wp-content/uploads/2023/07/data-exchange-files.png 1200w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /></figure>



<h2 class="wp-block-heading"><span id="data-storage-already-handles-this-%e2%80%93-and-it-does-it-better"><strong>Data storage already handles this – and it does it better</strong></span></h2>



<p class="wp-block-paragraph">These are only some of the issues that we encountered. There are more of them, some not even mentioned here, and others are yet to be discovered. To resolve them, we had to invest time in development, refactoring, and as a result, we now have a code base to maintain.&nbsp;</p>



<p class="wp-block-paragraph">But <strong>these problems were not unique to us.</strong> They were common problems that most data storage software handles out-of-the-box and does it better!</p>



<p class="wp-block-paragraph">The selection of software depends on the nature of the data. In our case, the data were events, and the producer application performed some batch processing on it. A more&nbsp;suitable storage option than the one we have been using <strong>would be a messaging system:</strong></p>



<p class="wp-block-paragraph">· Producers and consumers would write and read from a topic or a queue.</p>



<p class="wp-block-paragraph">· Data would be stored and replicated to another machine, also enabling Geo-redundancy.</p>



<p class="wp-block-paragraph">· Retention would be managed by the platform, eliminating the need for manual scripts.</p>



<p class="wp-block-paragraph">· Messaging systems have a&nbsp;<em>commit</em>&nbsp;feature to indicate that a record (or batch) has been processed.</p>



<p class="wp-block-paragraph">· Some messaging systems support consumer groups, allowing concurrent processing out-of-the-box.</p>



<p class="wp-block-paragraph">· Different consumers (applications) can be subscribed to the same topic.</p>



<p class="wp-block-paragraph">· Applications can be written in the most popular programming languages.</p>



<p class="wp-block-paragraph">· Monitoring tools for popular systems are already available.</p>



<p class="wp-block-paragraph">· Popular systems are usually well-documented.</p>



<p class="wp-block-paragraph">For other types of data, the solution might be a more suitable data store like <strong>Redis or Postgres</strong>. It all depends on what the data is about and how it can be presented.</p>



<p class="wp-block-paragraph"><strong>Exchanging data via files should be considered an anti-pattern</strong> and only be used when other approaches cannot be implemented (like when integrating with a legacy system).&nbsp;Each time some data is exchanged using files, a red flag should pop up, and you should ask yourself &#8211; what is a better way to do this? There probably is.</p>
<p>The post <a href="https://shiftmag.dev/from-file-frustration-to-streamlined-data-exchange-999/">From file frustration to streamlined data exchange: Rethinking the approach</a> appeared first on <a href="https://shiftmag.dev">ShiftMag</a>.</p>
]]></content:encoded>
					
		
		
			</item>
	</channel>
</rss>

<!--
Performance optimized by W3 Total Cache. Learn more: https://www.boldgrid.com/w3-total-cache/?utm_source=w3tc&utm_medium=footer_comment&utm_campaign=free_plugin

Page Caching using Disk: Enhanced 

Served from: shiftmag.dev @ 2026-08-23 04:38:48 by W3 Total Cache
-->