RAG Chunking Best-Practices Guide

Drawing on benchmarks and industry practice from multiple organizations, this guide presents default RAG chunking settings, parameter-tuning methods, and strategies for different document types.

March 18, 2026 · 10 min · 2050 words · Andy SI
Read more

Agent = Model + Harness

LangChain’s breakdown of the agent harness connects context engineering, memory, MCP, and the agent loop into one coherent map.

March 18, 2026 · 6 min · 1167 words · Andy SI
Read more

The Complete Guide to Concurrency and Parallelism in Python: The Evolution from One Thread to Multiple Cores

Starting with the fundamental distinction between concurrency and parallelism, this guide systematically explains when to use Python’s GIL, threading, multiprocessing, asyncio, and concurrent.futures, how they work together, and how to choose among them.

March 15, 2026 · 18 min · 3774 words · Andy SI
Read more

Python Async Programming: A Complete Guide from Iteration Protocols to Coroutines, Tasks, and Thread Pools

Build a practical mental model for Python async programming in web backends and AI applications, from iteration protocols, generators, and coroutine objects through awaitables, Tasks, Futures, the event loop, structured concurrency, async iteration, and thread and process pools.

March 15, 2026 · 25 min · 5122 words · Andy SI
Read more

The Complete Best-Practices Guide to Prompt Engineering

A long-form Prompt Engineering guide for AI application engineers, covering foundational principles, context design, task chains, injection defenses, agent prompt design, and evaluation-driven development.

March 11, 2026 · 19 min · 4045 words · Andy SI
Read more

Notes on Anthropic's Contextual Retrieval

Starting from Anthropic’s Contextual Retrieval article and Appendix II, these notes summarize the core method, experimental findings, and RAG architecture principles suitable for production.

March 11, 2026 · 5 min · 994 words · Andy SI
Read more

Essential Reading for Contextual Retrieval and RAG

A curated list of ten high-value articles on Contextual Retrieval, Context Engineering, and RAG evaluation, spanning Anthropic’s original publications, open-source implementation guides, and a review of 2025 trends.

March 9, 2026 · 4 min · 801 words · Andy SI
Read more

Apple M5 vs. M4: A Practical Comparison for AI Engineers

A practical breakdown for AI engineers of what the M5’s changes over the M4—from CPU, cache, and memory bandwidth to Neural Accelerators—mean for local LLM and diffusion inference.

March 6, 2026 · 13 min · 2757 words · Andy SI
Read more

What Is an LLM API KV Cache?

KV Cache is a key concept connecting Transformer theory with LLM engineering and deployment. Understanding it completes the path from how a model computes to how it runs.

March 5, 2026 · 17 min · 3558 words · Andy SI
Read more

The Complete Guide to LLM Chain-of-Thought (CoT)

Understand what LLM Chain-of-Thought (CoT) is and how prompt engineering can elicit Chain-of-Thought (CoT) from an LLM.

March 4, 2026 · 12 min · 2430 words · Andy SI
Read more