PUBLIC LESSON PREVIEW / LLM Applications and Retrieval
Map the application and budget its context
Read this sample without an account. Sign in for the full workspace, saved progress and checkpoints. This preview does not save activity or award credit.
Start with the system around the model
An LLM application is a collection of components, not a single clever prompt. A user asks a question, application code checks the request, a retriever selects information, a model or deterministic responder drafts an answer, and validators decide whether that answer can be shown. Logs and evaluation connect the behavior back to evidence. Each component can fail differently, so draw the flow before choosing tools. Retrieval-augmented generation adds relevant external material to the input used for answering. Retrieval does not update the model’s learned weights. It supplies information for the current request. A document assistant might retrieve a fictional handbook section about equipment returns, then answer from that evidence. If the source is missing or outdated, a fluent model cannot turn it into reliable evidence. The application must handle that absence explicitly.
Tokens are a model-specific representation
A tokenizer converts text into units that the model processes. Tokens may correspond to words, parts of words, punctuation, or other pieces, depending on the tokenizer and language. Counting whitespace-separated words is not an exact token count. Code, unusual identifiers, multilingual text, and repeated symbols can produce very different token-to-word relationships. Use the actual tokenizer or documented service accounting when accuracy matters. A context window constrains the material a model can consider for a request under its specific interface and limits. Instructions, conversation history, retrieved passages, tool results, and generated output compete for finite capacity in ways that depend on the model. Do not hard-code a model’s advertised limit into a general lesson. Check the current model documentation before integration and preserve a conservative margin for framing and unexpected input growth.
Budget evidence instead of dumping everything
More text is not automatically better evidence. A large bundle of loosely related passages can bury the one sentence needed to answer the question. Select concise, relevant chunks, remove redundant material, and preserve identifiers that let the answer cite its sources. Keep enough surrounding context to interpret qualifiers, dates, exceptions, and units. Truncating a return policy before its exception can reverse its meaning. Create a budget with separate allowances for instructions, the user question, essential history, evidence, output, and a safety margin. If evidence exceeds the allocation, rank and select it deliberately. Do not silently cut the last characters of a long document and hope the result stays coherent. The worked example uses declared token estimates solely to teach arithmetic; it does not tokenize real text or claim accurate provider usage.
Build an offline first slice
The first version of your course assistant needs no live model. Use a small corpus of invented handbook paragraphs, a basic retriever, and an extractive responder that quotes a relevant sentence with its source identifier. This lets you test data flow, source tracking, budgeting, and abstention before introducing probabilistic generation. Label the responder honestly so users do not confuse a deterministic teaching substitute with a language model. Write down the boundaries of responsibility. The retriever finds candidates, the answerer uses selected evidence, and application code enforces permissions and output contracts. A better prompt may improve phrasing but cannot repair a broken permission filter. Similarly, a larger context window cannot supply a missing policy update. Learning to locate the failing stage is the foundation for building useful LLM systems without treating every defect as a prompting problem.
Reserve context capacity before selecting evidence
The arithmetic leaves 2,200 estimated tokens for evidence, admits H1 and H2, and skips H3. A real application requires model-specific token accounting and relevance-aware ordering.
budget = 4000 # Illustrative capacity, not a named model's limit.
reserved = {"instructions": 400, "question": 120, "history": 300,
"output": 800, "margin": 180}
evidence_budget = budget - sum(reserved.values())
chunks = [("H1", 700), ("H2", 650), ("H3", 900)]
selected = []
used = 0
for chunk_id, estimated_tokens in chunks:
if used + estimated_tokens <= evidence_budget:
selected.append(chunk_id)
used += estimated_tokens
print("evidence_budget", evidence_budget)
print("selected", selected, "used", used)Ready to try it yourself?
ChatGPT sign-in takes you to OpenAI and back to CodeTrail. It keeps your learning account separate from other learners.
Open the full lesson ↗