<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Backend | 1byteone - Java Backend / AI Application Developer</title><link>https://1byteone.github.io/en/tags/backend/</link><atom:link href="https://1byteone.github.io/en/tags/backend/index.xml" rel="self" type="application/rss+xml"/><description>Backend</description><generator>HugoBlox Kit (https://hugoblox.com)</generator><language>en-us</language><lastBuildDate>Fri, 21 Aug 2026 00:00:00 +0000</lastBuildDate><image><url>https://1byteone.github.io/media/icon_hu_1c0e9cb08cfb822a.png</url><title>Backend</title><link>https://1byteone.github.io/en/tags/backend/</link></image><item><title>FastAPI Project Architecture: Layering Routes and Domain Services</title><link>https://1byteone.github.io/en/blog/fastapi-project-architecture/</link><pubDate>Fri, 21 Aug 2026 00:00:00 +0000</pubDate><guid>https://1byteone.github.io/en/blog/fastapi-project-architecture/</guid><description>&lt;p&gt;&lt;em&gt;Whiteboard: FastAPI — FastAPI Project Architecture.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;FastAPI makes it easy to start in one file—and easy to end up with an unmaintainable global script. A durable structure lets routes handle HTTP, services handle use cases, and repositories handle persistence.&lt;/p&gt;
&lt;h2 id="core-mental-model"&gt;Core mental model&lt;/h2&gt;
&lt;p&gt;A useful direction is router → service → repository/provider, with dependencies injected through function parameters. Domain services should not depend on Request or HTTP status codes, so they can be reused by jobs and tests.&lt;/p&gt;
&lt;h2 id="key-mechanics"&gt;Key mechanics&lt;/h2&gt;
&lt;p&gt;Load and validate configuration at startup. Create and close pools in the lifespan context. Map exceptions centrally to a stable error schema.&lt;/p&gt;
&lt;h2 id="python-example"&gt;Python example&lt;/h2&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-python" data-lang="python"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="nn"&gt;fastapi&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;APIRouter&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;Depends&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="nn"&gt;pydantic&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;BaseModel&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="n"&gt;router&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;APIRouter&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;prefix&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;&amp;#34;/v1/chat&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;tags&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;&amp;#34;chat&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;ChatRequest&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;BaseModel&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;ChatResponse&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;BaseModel&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="n"&gt;answer&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nd"&gt;@router.post&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;&amp;#34;&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;response_model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;ChatResponse&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;chat&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;req&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;ChatRequest&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;service&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;Depends&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;get_chat_service&lt;/span&gt;&lt;span class="p"&gt;)):&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;ChatResponse&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;answer&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;service&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;answer&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;req&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;h2 id="course-focus"&gt;Course focus&lt;/h2&gt;
&lt;p&gt;This article turns the whiteboard into explicit engineering boundaries: define inputs and outputs first, then decide how state, failures, and observability work. The examples use Python and focus on durable design principles rather than a particular provider version; verify APIs against the documentation for your installed dependencies.&lt;/p&gt;
&lt;h2 id="engineering-practice"&gt;Engineering practice&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Keep business rules in the application layer instead of hiding them in untestable prompts or route handlers.&lt;/li&gt;
&lt;li&gt;Add timeouts, bounded retries, and idempotency keys to external calls; retries are not a complete error strategy.&lt;/li&gt;
&lt;li&gt;Record a request id, latency, input version, model/index version, and outcome without logging sensitive raw content.&lt;/li&gt;
&lt;li&gt;Use a small fixed regression set first, then monitor quality and cost with sampled production traffic.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="common-mistakes"&gt;Common mistakes&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Drawing only the happy path and omitting timeouts, empty results, rate limits, and rollback paths.&lt;/li&gt;
&lt;li&gt;Letting one function parse input, call providers, build prompts, and persist data.&lt;/li&gt;
&lt;li&gt;Replacing typed contracts with string conventions that can only be verified by manual integration.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="production-checklist"&gt;Production checklist&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Inputs, outputs, and error responses have explicit schemas&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; External dependencies have timeouts, bounded retries, rate limits, and fallbacks&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Logs, metrics, and traces can be correlated to one request&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Critical paths have unit tests, integration tests, and offline evaluation samples&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Secrets, user content, and provider responses follow least-privilege and privacy rules&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="practice"&gt;Practice&lt;/h2&gt;
&lt;p&gt;Implement the smallest loop shown on the whiteboard. Inject a timeout, an empty result, and a malformed payload, then check whether the system remains stable and diagnosable. Add one metric that proves your optimization improved quality or latency.&lt;/p&gt;
&lt;h2 id="hands-on-exercise"&gt;Hands-on exercise&lt;/h2&gt;
&lt;p&gt;Split a model-calling route into api, services, providers, and schemas. Inject a fake provider into the service and use TestClient to verify error mapping.&lt;/p&gt;
&lt;h2 id="conclusion"&gt;Conclusion&lt;/h2&gt;
&lt;p&gt;An AI feature becomes maintainable when every arrow on the whiteboard maps to an input, an output, and a failure strategy.&lt;/p&gt;</description></item><item><title>Java Core Architecture</title><link>https://1byteone.github.io/en/blog/java-core-architecture/</link><pubDate>Fri, 21 Aug 2026 00:00:00 +0000</pubDate><guid>https://1byteone.github.io/en/blog/java-core-architecture/</guid><description>&lt;p&gt;&lt;em&gt;Whiteboard note: the goal is not to memorize boxes, but to connect each arrow to a request, a contract, and an operational signal.&lt;/em&gt;&lt;/p&gt;
&lt;h2 id="start-with-a-real-scenario"&gt;Start with a real scenario&lt;/h2&gt;
&lt;p&gt;Use this situation as the mental model: &lt;strong&gt;A payment service starts slowly, then becomes faster as JIT compiles its hot methods.&lt;/strong&gt;. A useful backend diagram answers four questions: where does input enter, who owns state, which boundary can fail, and how do we know the system is healthy? The flow is &lt;strong&gt;Source → Bytecode → JVM → Machine Code&lt;/strong&gt;.&lt;/p&gt;
&lt;h2 id="deconstructing-the-architecture"&gt;Deconstructing the architecture&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Contract&lt;/strong&gt;: define input, output, and error shape before adding implementation detail.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Mechanism&lt;/strong&gt;: Class Loader → load / link / initialize; Heap → objects + GC managed memory. This is the part that determines latency, throughput, and testability.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;State and resources&lt;/strong&gt;: Stack → frames + local variables; Method Area → class metadata. Explain creation, ownership, cleanup, and recovery.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Failure path&lt;/strong&gt;: PC Register + Native Method Stack; combine it with &lt;strong&gt;Execution Engine → Interpreter + JIT&lt;/strong&gt; when choosing a fallback, retry, or rollback.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="a-small-implementation-boundary"&gt;A small implementation boundary&lt;/h2&gt;
&lt;p&gt;The following sketch is intentionally small. It shows where a production implementation should place validation and ownership rather than pretending that a happy path is enough.&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;request → bounded resource → validated boundary → observable result
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Keep external systems behind an adapter, repository, client, or gateway. The domain layer should depend on a stable contract so tests can use fakes and incidents can be isolated to one boundary.&lt;/p&gt;
&lt;h2 id="engineering-decisions-for-the-scenario"&gt;Engineering decisions for the scenario&lt;/h2&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Protect dependencies first.&lt;/strong&gt; Pools, queues, semaphores, caches, and worker counts need explicit limits. An unbounded queue only hides overload.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Preserve correctness.&lt;/strong&gt; Use an idempotency key, transaction boundary, lock, version check, or schema validation where retries or concurrency can duplicate work.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Optimize with evidence.&lt;/strong&gt; Use traces, slow-query data, GC pauses, queue depth, or p99 latency before choosing a cache, index, batch size, or concurrency setting.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Make recovery testable.&lt;/strong&gt; Timeout, retry, rate limit, circuit breaking, rollback, and alerting should be exercised in a drill.&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id="common-mistakes"&gt;Common mistakes&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Drawing only the happy path while omitting timeout, empty data, rejection, rollback, or replica lag.&lt;/li&gt;
&lt;li&gt;Treating framework defaults as business contracts and discovering their limits after an upgrade.&lt;/li&gt;
&lt;li&gt;Scaling machines before measuring pool exhaustion, lock contention, slow SQL, allocation, or event-loop blocking.&lt;/li&gt;
&lt;li&gt;Treating logs as the only observability tool; metrics, traces, sampling, and redaction are also required.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="production-checklist"&gt;Production checklist&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Input, output, and error responses have explicit schemas&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Dependencies have timeout, retry budgets, rate limits, and fallback behavior&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Pools and queues expose capacity, waiting, and rejected-work metrics&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Important state has idempotency, transaction, or recovery semantics&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Logs, metrics, and traces share a request or correlation ID&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Regression tests, load tests, and failure drills use production-like data&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;
&lt;p&gt;You understand this whiteboard when you can start from one request, explain every arrow&amp;rsquo;s contract, state how each risk is handled, and name the signal that would tell you to change the design.&lt;/p&gt;</description></item><item><title>OpenAI API Fundamentals: Requests, Messages, Parameters, and Reliability</title><link>https://1byteone.github.io/en/blog/openai-api-fundamentals/</link><pubDate>Fri, 21 Aug 2026 00:00:00 +0000</pubDate><guid>https://1byteone.github.io/en/blog/openai-api-fundamentals/</guid><description>&lt;p&gt;&lt;em&gt;Whiteboard: OpenAI — OpenAI API Fundamentals.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;The core of model API integration is not memorizing an endpoint. Treat the network call as an unreliable dependency and design request contracts, timeouts, error classes, and usage tracking from day one.&lt;/p&gt;
&lt;h2 id="core-mental-model"&gt;Core mental model&lt;/h2&gt;
&lt;p&gt;A request contains system/developer rules, user input, and optional tools. The model output is a candidate result; the application must still validate, filter, and persist it.&lt;/p&gt;
&lt;h2 id="key-mechanics"&gt;Key mechanics&lt;/h2&gt;
&lt;p&gt;Separate 4xx input/auth errors, 429 rate limits, 5xx upstream failures, and local timeouts. Retry only transient classes with jittered exponential backoff and a total budget.&lt;/p&gt;
&lt;h2 id="python-example"&gt;Python example&lt;/h2&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-python" data-lang="python"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="nn"&gt;asyncio&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="nn"&gt;random&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;call_with_retry&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;attempts&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;n&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nb"&gt;range&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;attempts&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="k"&gt;try&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;responses&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;timeout&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;20&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="k"&gt;except&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="ne"&gt;TimeoutError&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;RateLimitError&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;n&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="n"&gt;attempts&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="k"&gt;raise&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;asyncio&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;sleep&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt; &lt;span class="o"&gt;**&lt;/span&gt; &lt;span class="n"&gt;n&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;random&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;random&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;h2 id="course-focus"&gt;Course focus&lt;/h2&gt;
&lt;p&gt;This article turns the whiteboard into explicit engineering boundaries: define inputs and outputs first, then decide how state, failures, and observability work. The examples use Python and focus on durable design principles rather than a particular provider version; verify APIs against the documentation for your installed dependencies.&lt;/p&gt;
&lt;h2 id="engineering-practice"&gt;Engineering practice&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Keep business rules in the application layer instead of hiding them in untestable prompts or route handlers.&lt;/li&gt;
&lt;li&gt;Add timeouts, bounded retries, and idempotency keys to external calls; retries are not a complete error strategy.&lt;/li&gt;
&lt;li&gt;Record a request id, latency, input version, model/index version, and outcome without logging sensitive raw content.&lt;/li&gt;
&lt;li&gt;Use a small fixed regression set first, then monitor quality and cost with sampled production traffic.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="common-mistakes"&gt;Common mistakes&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Drawing only the happy path and omitting timeouts, empty results, rate limits, and rollback paths.&lt;/li&gt;
&lt;li&gt;Letting one function parse input, call providers, build prompts, and persist data.&lt;/li&gt;
&lt;li&gt;Replacing typed contracts with string conventions that can only be verified by manual integration.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="production-checklist"&gt;Production checklist&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Inputs, outputs, and error responses have explicit schemas&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; External dependencies have timeouts, bounded retries, rate limits, and fallbacks&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Logs, metrics, and traces can be correlated to one request&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Critical paths have unit tests, integration tests, and offline evaluation samples&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Secrets, user content, and provider responses follow least-privilege and privacy rules&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="practice"&gt;Practice&lt;/h2&gt;
&lt;p&gt;Implement the smallest loop shown on the whiteboard. Inject a timeout, an empty result, and a malformed payload, then check whether the system remains stable and diagnosable. Add one metric that proves your optimization improved quality or latency.&lt;/p&gt;
&lt;h2 id="hands-on-exercise"&gt;Hands-on exercise&lt;/h2&gt;
&lt;p&gt;Add request id, latency, input/output tokens, and error class logging. Use a fake client that emits 429 and 400 to verify only 429 is retried.&lt;/p&gt;
&lt;h2 id="conclusion"&gt;Conclusion&lt;/h2&gt;
&lt;p&gt;An AI feature becomes maintainable when every arrow on the whiteboard maps to an input, an output, and a failure strategy.&lt;/p&gt;</description></item></channel></rss>