PUBLIC LESSON PREVIEW / Machine Learning Foundations
Frame the prediction and protect the split
Read this sample without an account. Sign in for the full workspace, saved progress and checkpoints. This preview does not save activity or award credit.
A prediction is a timed decision
A model starts with a question that can be checked later. Imagine a small parcel service deciding which shipments deserve an additional dispatch review. The prediction unit is one parcel at dispatch. The target is whether delivery occurs after its promised deadline. The decision is whether a dispatcher inspects the shipment. These are different objects: a predicted probability is not the eventual action, and neither is the ground truth. Write all three before touching an algorithm. Next draw a timeline. Package weight, route distance, and weather observed before dispatch might be available. The eventual delivery scan and a customer complaint submitted afterward are unavailable at decision time. A column can look ordinary while revealing the answer. Ask who created each field, when it became available, and whether historical values were overwritten. This creates a practical data contract rather than a list of convenient columns.
Give each partition a job
Training data teaches model parameters. Validation data helps choose features, algorithms, settings, and decision thresholds. Test data supports a final estimate after those decisions are frozen. A familiar allocation is roughly sixty, twenty, and twenty percent, but no ratio is universally correct. Rare outcomes, short histories, and expensive labels can require different plans. The purpose is independence, not a magic fraction. Write the split before exploring outcome relationships. General schema checks can inspect every file, but data-dependent choices must respect the experiment boundary. A single look at a test score is not automatically disastrous; repeated changes motivated by that score make it part of development. If that happens, say so and obtain a new untouched evaluation sample when possible. Renaming the same rows does not restore their independence.
Match the split to future use
A random split is useful when records are reasonably exchangeable and the intended future resembles the observed population. It is risky when several rows describe the same underlying person, device, household, or event. A grouped split keeps an entity entirely in one partition. Otherwise, a model may recognize an entity rather than learn a pattern that transfers to new entities. Temporal splits answer a different question: how well can older records predict later ones? Train on earlier periods and evaluate on later periods without shuffling. Leave a gap when labels mature slowly or feature windows overlap the boundary. Stratification can preserve approximate class proportions in random classification splits, but it does not solve time dependence or repeated entities. Select the split from the deployment question, then report what that split cannot establish.
Make the boundary inspectable
Store row identifiers alongside partition names so another learner can recreate the experiment. Assert that identifier sets do not overlap. For grouped data, also assert that group sets do not overlap. For temporal data, compare maximum training timestamps with minimum evaluation timestamps and inspect the label horizon. A seeded shuffle is convenient for demonstrations, but a seed alone cannot prove that the split is appropriate. The worked example partitions synthetic parcel identifiers chronologically. It deliberately has no fitted model: methodological decisions come first. Add a short experiment note describing the unit, target, decision time, label delay, permitted fields, split rule, and primary metric. This note should let a reviewer predict what evidence the eventual score will support. A well-defined modest experiment is more valuable than a spectacular number produced by an unclear boundary.
A chronological split manifest
Runs with Python alone and prints 60, 20, and 20 records. Chronology is illustrative; the synthetic outcome pattern is not evidence about real deliveries.
rows = [{"id": f"P{i:03d}", "day": i, "late": int(i % 7 == 0)}
for i in range(1, 101)]
train = [r for r in rows if r["day"] <= 60]
valid = [r for r in rows if 60 < r["day"] <= 80]
test = [r for r in rows if r["day"] > 80]
sets = [{r["id"] for r in part} for part in (train, valid, test)]
assert not (sets[0] & sets[1] or sets[0] & sets[2] or sets[1] & sets[2])
assert max(r["day"] for r in train) < min(r["day"] for r in valid)
print({"train": len(train), "validation": len(valid), "test": len(test)})Ready to try it yourself?
ChatGPT sign-in takes you to OpenAI and back to CodeTrail. It keeps your learning account separate from other learners.
Open the full lesson ↗