<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>FastAPI | 1byteone - Java Backend / AI Application Developer</title><link>https://1byteone.github.io/en/tags/fastapi/</link><atom:link href="https://1byteone.github.io/en/tags/fastapi/index.xml" rel="self" type="application/rss+xml"/><description>FastAPI</description><generator>HugoBlox Kit (https://hugoblox.com)</generator><language>en-us</language><lastBuildDate>Fri, 21 Aug 2026 00:00:00 +0000</lastBuildDate><image><url>https://1byteone.github.io/media/icon_hu_1c0e9cb08cfb822a.png</url><title>FastAPI</title><link>https://1byteone.github.io/en/tags/fastapi/</link></image><item><title>AI Engineering with Python</title><link>https://1byteone.github.io/en/blog/python-ai-engineering/</link><pubDate>Fri, 21 Aug 2026 00:00:00 +0000</pubDate><guid>https://1byteone.github.io/en/blog/python-ai-engineering/</guid><description>&lt;p&gt;&lt;em&gt;Whiteboard note: the goal is not to memorize boxes, but to connect each arrow to a request, a contract, and an operational signal.&lt;/em&gt;&lt;/p&gt;
&lt;h2 id="start-with-a-real-scenario"&gt;Start with a real scenario&lt;/h2&gt;
&lt;p&gt;Use this situation as the mental model: &lt;strong&gt;A support copilot answers only from authorized policy passages and returns a review state when evidence is weak.&lt;/strong&gt;. A useful backend diagram answers four questions: where does input enter, who owns state, which boundary can fail, and how do we know the system is healthy? The flow is &lt;strong&gt;Application → Prompt / RAG / Tool → Model → Validated Result&lt;/strong&gt;.&lt;/p&gt;
&lt;h2 id="deconstructing-the-architecture"&gt;Deconstructing the architecture&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Contract&lt;/strong&gt;: define input, output, and error shape before adding implementation detail.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Mechanism&lt;/strong&gt;: FastAPI → typed request / response; Retriever → evidence + ACL filter. This is the part that determines latency, throughput, and testability.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;State and resources&lt;/strong&gt;: Model call → timeout + retry budget; Tool → allowlist + argument validation. Explain creation, ownership, cleanup, and recovery.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Failure path&lt;/strong&gt;: Structured output → schema parse; combine it with &lt;strong&gt;Trace → prompt version + cost + quality&lt;/strong&gt; when choosing a fallback, retry, or rollback.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="a-small-implementation-boundary"&gt;A small implementation boundary&lt;/h2&gt;
&lt;p&gt;The following sketch is intentionally small. It shows where a production implementation should place validation and ownership rather than pretending that a happy path is enough.&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;request → bounded resource → validated boundary → observable result
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Keep external systems behind an adapter, repository, client, or gateway. The domain layer should depend on a stable contract so tests can use fakes and incidents can be isolated to one boundary.&lt;/p&gt;
&lt;h2 id="engineering-decisions-for-the-scenario"&gt;Engineering decisions for the scenario&lt;/h2&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Protect dependencies first.&lt;/strong&gt; Pools, queues, semaphores, caches, and worker counts need explicit limits. An unbounded queue only hides overload.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Preserve correctness.&lt;/strong&gt; Use an idempotency key, transaction boundary, lock, version check, or schema validation where retries or concurrency can duplicate work.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Optimize with evidence.&lt;/strong&gt; Use traces, slow-query data, GC pauses, queue depth, or p99 latency before choosing a cache, index, batch size, or concurrency setting.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Make recovery testable.&lt;/strong&gt; Timeout, retry, rate limit, circuit breaking, rollback, and alerting should be exercised in a drill.&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id="common-mistakes"&gt;Common mistakes&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Drawing only the happy path while omitting timeout, empty data, rejection, rollback, or replica lag.&lt;/li&gt;
&lt;li&gt;Treating framework defaults as business contracts and discovering their limits after an upgrade.&lt;/li&gt;
&lt;li&gt;Scaling machines before measuring pool exhaustion, lock contention, slow SQL, allocation, or event-loop blocking.&lt;/li&gt;
&lt;li&gt;Treating logs as the only observability tool; metrics, traces, sampling, and redaction are also required.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="production-checklist"&gt;Production checklist&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Input, output, and error responses have explicit schemas&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Dependencies have timeout, retry budgets, rate limits, and fallback behavior&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Pools and queues expose capacity, waiting, and rejected-work metrics&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Important state has idempotency, transaction, or recovery semantics&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Logs, metrics, and traces share a request or correlation ID&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Regression tests, load tests, and failure drills use production-like data&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;
&lt;p&gt;You understand this whiteboard when you can start from one request, explain every arrow&amp;rsquo;s contract, state how each risk is handled, and name the signal that would tell you to change the design.&lt;/p&gt;</description></item><item><title>FastAPI AI Streaming APIs with Server-Sent Events</title><link>https://1byteone.github.io/en/blog/fastapi-ai-streaming-api/</link><pubDate>Fri, 21 Aug 2026 00:00:00 +0000</pubDate><guid>https://1byteone.github.io/en/blog/fastapi-ai-streaming-api/</guid><description>&lt;p&gt;&lt;em&gt;Whiteboard: FastAPI — FastAPI AI Streaming APIs with Server-Sent Events.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;Streaming improves time-to-first-byte and perceived responsiveness; it does not make model computation faster. Define event framing, completion signals, error events, and cleanup on disconnect.&lt;/p&gt;
&lt;h2 id="core-mental-model"&gt;Core mental model&lt;/h2&gt;
&lt;p&gt;The server emits data events and the client parses event boundaries. Do not concatenate arbitrary unescaped text into SSE; prefer JSON events with a stable type field.&lt;/p&gt;
&lt;h2 id="key-mechanics"&gt;Key mechanics&lt;/h2&gt;
&lt;p&gt;Use try/finally in the generator to release upstream resources. Include a request id in each event. Convert model failures into an error event and close with an explicit done signal.&lt;/p&gt;
&lt;h2 id="python-example"&gt;Python example&lt;/h2&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-python" data-lang="python"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="nn"&gt;json&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="nn"&gt;fastapi.responses&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;StreamingResponse&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;events&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="k"&gt;try&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;token&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;provider&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;stream&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="k"&gt;yield&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="s2"&gt;&amp;#34;data: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;dumps&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;&lt;span class="s1"&gt;&amp;#39;type&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s1"&gt;&amp;#39;token&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;&amp;#39;text&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;token&lt;/span&gt;&lt;span class="p"&gt;})&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="se"&gt;\n\n&lt;/span&gt;&lt;span class="s2"&gt;&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="k"&gt;yield&lt;/span&gt; &lt;span class="s2"&gt;&amp;#34;data: {&lt;/span&gt;&lt;span class="se"&gt;\&amp;#34;&lt;/span&gt;&lt;span class="s2"&gt;type&lt;/span&gt;&lt;span class="se"&gt;\&amp;#34;&lt;/span&gt;&lt;span class="s2"&gt;:&lt;/span&gt;&lt;span class="se"&gt;\&amp;#34;&lt;/span&gt;&lt;span class="s2"&gt;done&lt;/span&gt;&lt;span class="se"&gt;\&amp;#34;&lt;/span&gt;&lt;span class="s2"&gt;}&lt;/span&gt;&lt;span class="se"&gt;\n\n&lt;/span&gt;&lt;span class="s2"&gt;&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="k"&gt;except&lt;/span&gt; &lt;span class="ne"&gt;Exception&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="k"&gt;yield&lt;/span&gt; &lt;span class="s2"&gt;&amp;#34;data: {&lt;/span&gt;&lt;span class="se"&gt;\&amp;#34;&lt;/span&gt;&lt;span class="s2"&gt;type&lt;/span&gt;&lt;span class="se"&gt;\&amp;#34;&lt;/span&gt;&lt;span class="s2"&gt;:&lt;/span&gt;&lt;span class="se"&gt;\&amp;#34;&lt;/span&gt;&lt;span class="s2"&gt;error&lt;/span&gt;&lt;span class="se"&gt;\&amp;#34;&lt;/span&gt;&lt;span class="s2"&gt;,&lt;/span&gt;&lt;span class="se"&gt;\&amp;#34;&lt;/span&gt;&lt;span class="s2"&gt;code&lt;/span&gt;&lt;span class="se"&gt;\&amp;#34;&lt;/span&gt;&lt;span class="s2"&gt;:&lt;/span&gt;&lt;span class="se"&gt;\&amp;#34;&lt;/span&gt;&lt;span class="s2"&gt;upstream_failed&lt;/span&gt;&lt;span class="se"&gt;\&amp;#34;&lt;/span&gt;&lt;span class="s2"&gt;}&lt;/span&gt;&lt;span class="se"&gt;\n\n&lt;/span&gt;&lt;span class="s2"&gt;&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nd"&gt;@app.post&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;&amp;#34;/v1/chat/stream&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;stream&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;req&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;ChatRequest&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;StreamingResponse&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;events&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;req&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="n"&gt;media_type&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;&amp;#34;text/event-stream&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;h2 id="course-focus"&gt;Course focus&lt;/h2&gt;
&lt;p&gt;This article turns the whiteboard into explicit engineering boundaries: define inputs and outputs first, then decide how state, failures, and observability work. The examples use Python and focus on durable design principles rather than a particular provider version; verify APIs against the documentation for your installed dependencies.&lt;/p&gt;
&lt;h2 id="engineering-practice"&gt;Engineering practice&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Keep business rules in the application layer instead of hiding them in untestable prompts or route handlers.&lt;/li&gt;
&lt;li&gt;Add timeouts, bounded retries, and idempotency keys to external calls; retries are not a complete error strategy.&lt;/li&gt;
&lt;li&gt;Record a request id, latency, input version, model/index version, and outcome without logging sensitive raw content.&lt;/li&gt;
&lt;li&gt;Use a small fixed regression set first, then monitor quality and cost with sampled production traffic.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="common-mistakes"&gt;Common mistakes&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Drawing only the happy path and omitting timeouts, empty results, rate limits, and rollback paths.&lt;/li&gt;
&lt;li&gt;Letting one function parse input, call providers, build prompts, and persist data.&lt;/li&gt;
&lt;li&gt;Replacing typed contracts with string conventions that can only be verified by manual integration.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="production-checklist"&gt;Production checklist&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Inputs, outputs, and error responses have explicit schemas&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; External dependencies have timeouts, bounded retries, rate limits, and fallbacks&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Logs, metrics, and traces can be correlated to one request&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Critical paths have unit tests, integration tests, and offline evaluation samples&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Secrets, user content, and provider responses follow least-privilege and privacy rules&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="practice"&gt;Practice&lt;/h2&gt;
&lt;p&gt;Implement the smallest loop shown on the whiteboard. Inject a timeout, an empty result, and a malformed payload, then check whether the system remains stable and diagnosable. Add one metric that proves your optimization improved quality or latency.&lt;/p&gt;
&lt;h2 id="hands-on-exercise"&gt;Hands-on exercise&lt;/h2&gt;
&lt;p&gt;Write a browser SSE client and test normal completion, upstream failure, refresh, and network disconnect. Track TTFT, full latency, and cancellation rate.&lt;/p&gt;
&lt;h2 id="conclusion"&gt;Conclusion&lt;/h2&gt;
&lt;p&gt;An AI feature becomes maintainable when every arrow on the whiteboard maps to an input, an output, and a failure strategy.&lt;/p&gt;</description></item><item><title>FastAPI Async and Dependency Injection: Keep the Event Loop Healthy</title><link>https://1byteone.github.io/en/blog/fastapi-async-dependency-injection/</link><pubDate>Fri, 21 Aug 2026 00:00:00 +0000</pubDate><guid>https://1byteone.github.io/en/blog/fastapi-async-dependency-injection/</guid><description>&lt;p&gt;&lt;em&gt;Whiteboard: FastAPI — FastAPI Async and Dependency Injection.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;async def does not make blocking code faster. File I/O, synchronous SDKs, and CPU-heavy work can stall the event loop and slow every request.&lt;/p&gt;
&lt;h2 id="core-mental-model"&gt;Core mental model&lt;/h2&gt;
&lt;p&gt;Async routes are appropriate for async network I/O. Isolate synchronous dependencies explicitly, and move CPU-heavy work to a queue or process pool. Dependency injection should make lifetimes visible and replacements easy.&lt;/p&gt;
&lt;h2 id="key-mechanics"&gt;Key mechanics&lt;/h2&gt;
&lt;p&gt;Create shared clients in lifespan and reuse their connections. Use a semaphore to bound concurrent provider calls so your service does not overwhelm an upstream.&lt;/p&gt;
&lt;h2 id="python-example"&gt;Python example&lt;/h2&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-python" data-lang="python"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="nn"&gt;asyncio&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="nn"&gt;contextlib&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;asynccontextmanager&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="nn"&gt;fastapi&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;FastAPI&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="n"&gt;limit&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;asyncio&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Semaphore&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;20&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;call_provider&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;payload&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="n"&gt;limit&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;post&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;&amp;#34;/responses&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;payload&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;timeout&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;15&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nd"&gt;@asynccontextmanager&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;lifespan&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;app&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;FastAPI&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="n"&gt;app&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;state&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;make_async_client&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="k"&gt;yield&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;app&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;state&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;aclose&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="n"&gt;app&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;FastAPI&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;lifespan&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;lifespan&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;h2 id="course-focus"&gt;Course focus&lt;/h2&gt;
&lt;p&gt;This article turns the whiteboard into explicit engineering boundaries: define inputs and outputs first, then decide how state, failures, and observability work. The examples use Python and focus on durable design principles rather than a particular provider version; verify APIs against the documentation for your installed dependencies.&lt;/p&gt;
&lt;h2 id="engineering-practice"&gt;Engineering practice&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Keep business rules in the application layer instead of hiding them in untestable prompts or route handlers.&lt;/li&gt;
&lt;li&gt;Add timeouts, bounded retries, and idempotency keys to external calls; retries are not a complete error strategy.&lt;/li&gt;
&lt;li&gt;Record a request id, latency, input version, model/index version, and outcome without logging sensitive raw content.&lt;/li&gt;
&lt;li&gt;Use a small fixed regression set first, then monitor quality and cost with sampled production traffic.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="common-mistakes"&gt;Common mistakes&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Drawing only the happy path and omitting timeouts, empty results, rate limits, and rollback paths.&lt;/li&gt;
&lt;li&gt;Letting one function parse input, call providers, build prompts, and persist data.&lt;/li&gt;
&lt;li&gt;Replacing typed contracts with string conventions that can only be verified by manual integration.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="production-checklist"&gt;Production checklist&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Inputs, outputs, and error responses have explicit schemas&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; External dependencies have timeouts, bounded retries, rate limits, and fallbacks&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Logs, metrics, and traces can be correlated to one request&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Critical paths have unit tests, integration tests, and offline evaluation samples&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Secrets, user content, and provider responses follow least-privilege and privacy rules&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="practice"&gt;Practice&lt;/h2&gt;
&lt;p&gt;Implement the smallest loop shown on the whiteboard. Inject a timeout, an empty result, and a malformed payload, then check whether the system remains stable and diagnosable. Add one metric that proves your optimization improved quality or latency.&lt;/p&gt;
&lt;h2 id="hands-on-exercise"&gt;Hands-on exercise&lt;/h2&gt;
&lt;p&gt;Benchmark 50 concurrent requests with a synchronous fake client and an async fake client. Add a semaphore and compare upstream error rate and mean latency.&lt;/p&gt;
&lt;h2 id="conclusion"&gt;Conclusion&lt;/h2&gt;
&lt;p&gt;An AI feature becomes maintainable when every arrow on the whiteboard maps to an input, an output, and a failure strategy.&lt;/p&gt;</description></item><item><title>FastAPI Project Architecture: Layering Routes and Domain Services</title><link>https://1byteone.github.io/en/blog/fastapi-project-architecture/</link><pubDate>Fri, 21 Aug 2026 00:00:00 +0000</pubDate><guid>https://1byteone.github.io/en/blog/fastapi-project-architecture/</guid><description>&lt;p&gt;&lt;em&gt;Whiteboard: FastAPI — FastAPI Project Architecture.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;FastAPI makes it easy to start in one file—and easy to end up with an unmaintainable global script. A durable structure lets routes handle HTTP, services handle use cases, and repositories handle persistence.&lt;/p&gt;
&lt;h2 id="core-mental-model"&gt;Core mental model&lt;/h2&gt;
&lt;p&gt;A useful direction is router → service → repository/provider, with dependencies injected through function parameters. Domain services should not depend on Request or HTTP status codes, so they can be reused by jobs and tests.&lt;/p&gt;
&lt;h2 id="key-mechanics"&gt;Key mechanics&lt;/h2&gt;
&lt;p&gt;Load and validate configuration at startup. Create and close pools in the lifespan context. Map exceptions centrally to a stable error schema.&lt;/p&gt;
&lt;h2 id="python-example"&gt;Python example&lt;/h2&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-python" data-lang="python"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="nn"&gt;fastapi&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;APIRouter&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;Depends&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="nn"&gt;pydantic&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;BaseModel&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="n"&gt;router&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;APIRouter&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;prefix&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;&amp;#34;/v1/chat&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;tags&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;&amp;#34;chat&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;ChatRequest&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;BaseModel&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;ChatResponse&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;BaseModel&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="n"&gt;answer&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nd"&gt;@router.post&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;&amp;#34;&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;response_model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;ChatResponse&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;chat&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;req&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;ChatRequest&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;service&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;Depends&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;get_chat_service&lt;/span&gt;&lt;span class="p"&gt;)):&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;ChatResponse&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;answer&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;service&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;answer&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;req&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;h2 id="course-focus"&gt;Course focus&lt;/h2&gt;
&lt;p&gt;This article turns the whiteboard into explicit engineering boundaries: define inputs and outputs first, then decide how state, failures, and observability work. The examples use Python and focus on durable design principles rather than a particular provider version; verify APIs against the documentation for your installed dependencies.&lt;/p&gt;
&lt;h2 id="engineering-practice"&gt;Engineering practice&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Keep business rules in the application layer instead of hiding them in untestable prompts or route handlers.&lt;/li&gt;
&lt;li&gt;Add timeouts, bounded retries, and idempotency keys to external calls; retries are not a complete error strategy.&lt;/li&gt;
&lt;li&gt;Record a request id, latency, input version, model/index version, and outcome without logging sensitive raw content.&lt;/li&gt;
&lt;li&gt;Use a small fixed regression set first, then monitor quality and cost with sampled production traffic.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="common-mistakes"&gt;Common mistakes&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Drawing only the happy path and omitting timeouts, empty results, rate limits, and rollback paths.&lt;/li&gt;
&lt;li&gt;Letting one function parse input, call providers, build prompts, and persist data.&lt;/li&gt;
&lt;li&gt;Replacing typed contracts with string conventions that can only be verified by manual integration.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="production-checklist"&gt;Production checklist&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Inputs, outputs, and error responses have explicit schemas&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; External dependencies have timeouts, bounded retries, rate limits, and fallbacks&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Logs, metrics, and traces can be correlated to one request&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Critical paths have unit tests, integration tests, and offline evaluation samples&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Secrets, user content, and provider responses follow least-privilege and privacy rules&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="practice"&gt;Practice&lt;/h2&gt;
&lt;p&gt;Implement the smallest loop shown on the whiteboard. Inject a timeout, an empty result, and a malformed payload, then check whether the system remains stable and diagnosable. Add one metric that proves your optimization improved quality or latency.&lt;/p&gt;
&lt;h2 id="hands-on-exercise"&gt;Hands-on exercise&lt;/h2&gt;
&lt;p&gt;Split a model-calling route into api, services, providers, and schemas. Inject a fake provider into the service and use TestClient to verify error mapping.&lt;/p&gt;
&lt;h2 id="conclusion"&gt;Conclusion&lt;/h2&gt;
&lt;p&gt;An AI feature becomes maintainable when every arrow on the whiteboard maps to an input, an output, and a failure strategy.&lt;/p&gt;</description></item><item><title>FastAPI Request and Response Contracts with Pydantic</title><link>https://1byteone.github.io/en/blog/fastapi-request-response-pydantic/</link><pubDate>Fri, 21 Aug 2026 00:00:00 +0000</pubDate><guid>https://1byteone.github.io/en/blog/fastapi-request-response-pydantic/</guid><description>&lt;p&gt;&lt;em&gt;Whiteboard: FastAPI — FastAPI Request and Response Contracts with Pydantic.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;AI endpoints receive untrusted client input and often relay unstable model or third-party output. Pydantic schemas create the first boundary by turning data that looks right into data that has been validated.&lt;/p&gt;
&lt;h2 id="core-mental-model"&gt;Core mental model&lt;/h2&gt;
&lt;p&gt;Request models validate inputs, domain models enforce business invariants, and response models define the public surface. Do not return ORM objects or raw provider JSON directly.&lt;/p&gt;
&lt;h2 id="key-mechanics"&gt;Key mechanics&lt;/h2&gt;
&lt;p&gt;Bound string lengths, use Literal for enumerations, and model nested structures. Parse model output before returning it; route failures through an observable 502/422 branch.&lt;/p&gt;
&lt;h2 id="python-example"&gt;Python example&lt;/h2&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-python" data-lang="python"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="nn"&gt;typing&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Literal&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="nn"&gt;pydantic&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;BaseModel&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;Field&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;ChatRequest&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;BaseModel&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;Field&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;min_length&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;max_length&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;4000&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="n"&gt;mode&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;Literal&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;&amp;#34;answer&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;&amp;#34;summarize&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;&amp;#34;answer&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;ChatResponse&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;BaseModel&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="n"&gt;request_id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="n"&gt;answer&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="n"&gt;citations&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;list&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;Field&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;default_factory&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nb"&gt;list&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;h2 id="course-focus"&gt;Course focus&lt;/h2&gt;
&lt;p&gt;This article turns the whiteboard into explicit engineering boundaries: define inputs and outputs first, then decide how state, failures, and observability work. The examples use Python and focus on durable design principles rather than a particular provider version; verify APIs against the documentation for your installed dependencies.&lt;/p&gt;
&lt;h2 id="engineering-practice"&gt;Engineering practice&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Keep business rules in the application layer instead of hiding them in untestable prompts or route handlers.&lt;/li&gt;
&lt;li&gt;Add timeouts, bounded retries, and idempotency keys to external calls; retries are not a complete error strategy.&lt;/li&gt;
&lt;li&gt;Record a request id, latency, input version, model/index version, and outcome without logging sensitive raw content.&lt;/li&gt;
&lt;li&gt;Use a small fixed regression set first, then monitor quality and cost with sampled production traffic.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="common-mistakes"&gt;Common mistakes&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Drawing only the happy path and omitting timeouts, empty results, rate limits, and rollback paths.&lt;/li&gt;
&lt;li&gt;Letting one function parse input, call providers, build prompts, and persist data.&lt;/li&gt;
&lt;li&gt;Replacing typed contracts with string conventions that can only be verified by manual integration.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="production-checklist"&gt;Production checklist&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Inputs, outputs, and error responses have explicit schemas&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; External dependencies have timeouts, bounded retries, rate limits, and fallbacks&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Logs, metrics, and traces can be correlated to one request&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Critical paths have unit tests, integration tests, and offline evaluation samples&lt;/li&gt;
&lt;li&gt;&lt;input disabled="" type="checkbox"&gt; Secrets, user content, and provider responses follow least-privilege and privacy rules&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="practice"&gt;Practice&lt;/h2&gt;
&lt;p&gt;Implement the smallest loop shown on the whiteboard. Inject a timeout, an empty result, and a malformed payload, then check whether the system remains stable and diagnosable. Add one metric that proves your optimization improved quality or latency.&lt;/p&gt;
&lt;h2 id="hands-on-exercise"&gt;Hands-on exercise&lt;/h2&gt;
&lt;p&gt;Add request_id, length limits, a mode enum, and citations to a Q&amp;amp;A endpoint. Submit an empty string, oversized text, unknown mode, and missing fields and verify stable client errors.&lt;/p&gt;
&lt;h2 id="conclusion"&gt;Conclusion&lt;/h2&gt;
&lt;p&gt;An AI feature becomes maintainable when every arrow on the whiteboard maps to an input, an output, and a failure strategy.&lt;/p&gt;</description></item></channel></rss>