<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Production | 1byteone - Java Backend / AI Application Developer</title><link>https://1byteone.github.io/en/tags/production/</link><atom:link href="https://1byteone.github.io/en/tags/production/index.xml" rel="self" type="application/rss+xml"/><description>Production</description><generator>HugoBlox Kit (https://hugoblox.com)</generator><language>en-us</language><lastBuildDate>Fri, 21 Aug 2026 00:00:00 +0000</lastBuildDate><image><url>https://1byteone.github.io/media/icon_hu_1c0e9cb08cfb822a.png</url><title>Production</title><link>https://1byteone.github.io/en/tags/production/</link></image><item><title>Production RAG Architecture: Permissions, Versions, and Observability</title><link>https://1byteone.github.io/en/blog/rag-production-architecture/</link><pubDate>Fri, 21 Aug 2026 00:00:00 +0000</pubDate><guid>https://1byteone.github.io/en/blog/rag-production-architecture/</guid><description>&lt;p&gt;&lt;em&gt;Whiteboard: RAG — Production RAG Architecture.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;The gap between a RAG demo and production is mostly boundary conditions: who can see a document, how indexes update, how answers cite evidence, and what happens when an upstream is unavailable.&lt;/p&gt;
&lt;h2 id="core-mental-model"&gt;Core mental model&lt;/h2&gt;
&lt;p&gt;Separate ingestion, indexing, query serving, and evaluation planes. Online traffic reads only the active index while new versions build in the background and switch atomically.&lt;/p&gt;
&lt;h2 id="key-mechanics"&gt;Key mechanics&lt;/h2&gt;
&lt;p&gt;Enforce permissions during retrieval. Cache keys should include tenant, a permission digest, query-normalization version, and index version. Deletions must reach every replica.&lt;/p&gt;
&lt;h2 id="python-example"&gt;Python example&lt;/h2&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-python" data-lang="python"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="nn"&gt;dataclasses&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;dataclass&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nd"&gt;@dataclass&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;QueryContext&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="n"&gt;tenant_id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="n"&gt;allowed_collections&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;set&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="n"&gt;index_version&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;cache_key&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;QueryContext&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="n"&gt;scopes&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;&amp;#34;,&amp;#34;&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;join&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nb"&gt;sorted&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ctx&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;allowed_collections&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="s2"&gt;&amp;#34;rag:v2:&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;ctx&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;tenant_id&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;:&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;ctx&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;index_version&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;:&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;scopes&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;:&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;normalize&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;h2 id="course-focus"&gt;Course focus&lt;/h2&gt;
&lt;p&gt;This article turns the whiteboard into explicit engineering boundaries: define inputs and outputs first, then decide how state, failures, and observability work. The examples use Python and focus on durable design principles rather than a particular provider version; verify APIs against the documentation for your installed dependencies.&lt;/p&gt;
&lt;h2 id="engineering-practice"&gt;Engineering practice&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Keep business rules in the application layer instead of hiding them in untestable prompts or route handlers.&lt;/li&gt;
&lt;li&gt;Add timeouts, bounded retries, and idempotency keys to external calls; retries are not a complete error strategy.&lt;/li&gt;
&lt;li&gt;Record a request id, latency, input version, model/index version, and outcome without logging sensitive raw content.&lt;/li&gt;
&lt;li&gt;Use a small fixed regression set first, then monitor quality and cost with sampled production traffic.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="common-mistakes"&gt;Common mistakes&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Drawing only the happy path and omitting timeouts, empty results, rate limits, and rollback paths.&lt;/li&gt;
&lt;li&gt;Letting one function parse input, call providers, build prompts, and persist data.&lt;/li&gt;
&lt;li&gt;Replacing typed contracts with string conventions that can only be verified by manual integration.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="production-checklist"&gt;Production checklist&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Inputs, outputs, and error responses have explicit schemas&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; External dependencies have timeouts, bounded retries, rate limits, and fallbacks&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Logs, metrics, and traces can be correlated to one request&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Critical paths have unit tests, integration tests, and offline evaluation samples&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Secrets, user content, and provider responses follow least-privilege and privacy rules&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="practice"&gt;Practice&lt;/h2&gt;
&lt;p&gt;Implement the smallest loop shown on the whiteboard. Inject a timeout, an empty result, and a malformed payload, then check whether the system remains stable and diagnosable. Add one metric that proves your optimization improved quality or latency.&lt;/p&gt;
&lt;h2 id="hands-on-exercise"&gt;Hands-on exercise&lt;/h2&gt;
&lt;p&gt;Design active and candidate indexes. Simulate a failed publish and a document deletion, verify traffic never receives unauthorized stale data, and persist evidence ids for every answer.&lt;/p&gt;
&lt;h2 id="conclusion"&gt;Conclusion&lt;/h2&gt;
&lt;p&gt;An AI feature becomes maintainable when every arrow on the whiteboard maps to an input, an output, and a failure strategy.&lt;/p&gt;</description></item><item><title>Production Spring Boot Architecture</title><link>https://1byteone.github.io/en/blog/spring-boot-production-architecture/</link><pubDate>Fri, 21 Aug 2026 00:00:00 +0000</pubDate><guid>https://1byteone.github.io/en/blog/spring-boot-production-architecture/</guid><description>&lt;p&gt;&lt;em&gt;Whiteboard note: the goal is not to memorize boxes, but to connect each arrow to a request, a contract, and an operational signal.&lt;/em&gt;&lt;/p&gt;
&lt;h2 id="start-with-a-real-scenario"&gt;Start with a real scenario&lt;/h2&gt;
&lt;p&gt;Use this situation as the mental model: &lt;strong&gt;A flash-sale service protects the database with cache, bounded writes, idempotency, and asynchronous inventory events.&lt;/strong&gt;. A useful backend diagram answers four questions: where does input enter, who owns state, which boundary can fail, and how do we know the system is healthy? The flow is &lt;strong&gt;Gateway → Service → Cache / DB / MQ → Observability&lt;/strong&gt;.&lt;/p&gt;
&lt;h2 id="deconstructing-the-architecture"&gt;Deconstructing the architecture&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Contract&lt;/strong&gt;: define input, output, and error shape before adding implementation detail.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Mechanism&lt;/strong&gt;: Nginx / Gateway → auth + rate limit; Service → idempotency + transaction. This is the part that determines latency, throughput, and testability.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;State and resources&lt;/strong&gt;: Redis → cache-aside with TTL; MySQL → source of truth. Explain creation, ownership, cleanup, and recovery.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Failure path&lt;/strong&gt;: RocketMQ → async side effects; combine it with &lt;strong&gt;Metrics + Logs + Traces → feedback loop&lt;/strong&gt; when choosing a fallback, retry, or rollback.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="a-small-implementation-boundary"&gt;A small implementation boundary&lt;/h2&gt;
&lt;p&gt;The following sketch is intentionally small. It shows where a production implementation should place validation and ownership rather than pretending that a happy path is enough.&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;request → bounded resource → validated boundary → observable result
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Keep external systems behind an adapter, repository, client, or gateway. The domain layer should depend on a stable contract so tests can use fakes and incidents can be isolated to one boundary.&lt;/p&gt;
&lt;h2 id="engineering-decisions-for-the-scenario"&gt;Engineering decisions for the scenario&lt;/h2&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Protect dependencies first.&lt;/strong&gt; Pools, queues, semaphores, caches, and worker counts need explicit limits. An unbounded queue only hides overload.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Preserve correctness.&lt;/strong&gt; Use an idempotency key, transaction boundary, lock, version check, or schema validation where retries or concurrency can duplicate work.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Optimize with evidence.&lt;/strong&gt; Use traces, slow-query data, GC pauses, queue depth, or p99 latency before choosing a cache, index, batch size, or concurrency setting.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Make recovery testable.&lt;/strong&gt; Timeout, retry, rate limit, circuit breaking, rollback, and alerting should be exercised in a drill.&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id="common-mistakes"&gt;Common mistakes&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Drawing only the happy path while omitting timeout, empty data, rejection, rollback, or replica lag.&lt;/li&gt;
&lt;li&gt;Treating framework defaults as business contracts and discovering their limits after an upgrade.&lt;/li&gt;
&lt;li&gt;Scaling machines before measuring pool exhaustion, lock contention, slow SQL, allocation, or event-loop blocking.&lt;/li&gt;
&lt;li&gt;Treating logs as the only observability tool; metrics, traces, sampling, and redaction are also required.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="production-checklist"&gt;Production checklist&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Input, output, and error responses have explicit schemas&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Dependencies have timeout, retry budgets, rate limits, and fallback behavior&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Pools and queues expose capacity, waiting, and rejected-work metrics&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Important state has idempotency, transaction, or recovery semantics&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Logs, metrics, and traces share a request or correlation ID&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Regression tests, load tests, and failure drills use production-like data&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;
&lt;p&gt;You understand this whiteboard when you can start from one request, explain every arrow&amp;rsquo;s contract, state how each risk is handled, and name the signal that would tell you to change the design.&lt;/p&gt;</description></item></channel></rss>