ML And AI Basics
Core Terminology
| Term | Meaning |
|---|---|
| Feature | An input variable used to make a prediction |
| Label / target | The value a supervised model learns to predict |
| Model | A function (with learned parameters) mapping inputs to outputs |
| Parameters | Values learned during training (e.g. weights) |
| Hyperparameters | Values set before training (learning rate, layer count) — not learned |
| Inference | Running a trained model on new input to get a prediction |
| Dataset split | Training (fit), validation (tune), test (final unbiased check) |
Types Of Learning
- Supervised — learn a mapping from labeled input→output pairs (classification, regression).
- Unsupervised — find structure in unlabeled data (clustering, dimensionality reduction).
- Self-supervised — generate labels from the data itself (e.g. predict the next token) — how most large language models are pretrained.
- Reinforcement learning — an agent learns by taking actions in an environment and receiving reward signals (used in game-playing agents, RLHF fine-tuning).
Model Training Loop
for epoch in range(num_epochs):
for batch in train_loader:
predictions = model(batch.inputs)
loss = loss_fn(predictions, batch.labels)
loss.backward() # compute gradients (backpropagation)
optimizer.step() # update weights
optimizer.zero_grad() # reset gradients for next batch
validate(model, val_loader) # check generalization each epoch
Backpropagation computes how much each parameter contributed to the error via the chain rule; the optimizer (e.g. SGD, Adam) nudges parameters to reduce loss. An epoch is one full pass over the training data; a batch is a subset processed per gradient step.
Evaluation Metrics
| Task | Metrics |
|---|---|
| Classification | accuracy, precision, recall, F1, ROC-AUC, confusion matrix |
| Regression | MAE, MSE, RMSE, R² |
| Ranking / retrieval | precision@k, recall@k, NDCG, MRR |
| Generation (LLM) | perplexity, BLEU/ROUGE (translation/summarization), human/LLM-judged eval |
Precision = of predicted positives, how many were actually positive. Recall = of actual positives, how many were caught. F1 is their harmonic mean — use it when classes are imbalanced and accuracy alone is misleading.
Overfitting And Regularization
Overfitting: the model memorizes training data (including noise) instead of learning generalizable patterns — low training loss, high validation loss.
Mitigations:
- More/varied data — the most reliable fix.
- Regularization — L1/L2 weight penalties, dropout (randomly zero activations during training).
- Early stopping — stop training when validation loss stops improving.
- Simpler model / fewer parameters — reduce capacity to match the data available.
- Cross-validation — check generalization isn’t a fluke of one train/val split.
Underfitting is the opposite failure — the model is too simple to capture the pattern (high error on both train and validation).
Neural Network Basics
A neural network is layers of linear transformations (weights + bias) interleaved with nonlinear activation functions, trained end-to-end via gradient descent.
input → [linear → activation] × N layers → output
| Activation | Typical use |
|---|---|
| ReLU | hidden layers (default choice, cheap, avoids vanishing gradients) |
| Sigmoid | binary classification output (0-1 probability) |
| Softmax | multi-class classification output (probability distribution) |
| Tanh | occasionally in RNNs/older architectures |
Without nonlinear activations, stacking linear layers collapses to a single linear function — depth would add no expressive power.
Common Architectures
| Architecture | Good for |
|---|---|
| CNN (convolutional) | images, spatial/grid data — learns local patterns via shared filters |
| RNN / LSTM / GRU | sequences (legacy for NLP/time series) — processes step by step, struggles with long-range dependencies |
| Transformer | sequences, now dominant for NLP and increasingly vision — attention lets every position see every other position directly |
| Autoencoder | compression, anomaly detection, denoising — learns to reconstruct its input through a narrow bottleneck |
| GAN / diffusion | generative image/audio synthesis |
LLM Basics
Large language models are transformer-based, trained to predict the next token over massive text corpora (self-supervised pretraining), then adjusted via instruction tuning and RLHF/DPO to follow instructions and align with human preference.
- Token — a chunk of text (roughly a word-piece or subword) the model actually processes; not the same as a character or whole word.
- Context window — the maximum number of tokens a model can attend to in one request (input + output combined).
- Temperature — sampling randomness at inference: near 0 is deterministic/focused, higher values are more varied/creative.
- Hallucination — the model generates plausible-sounding but factually wrong or unsupported output — an inherent risk since generation is a probability model, not a lookup.
Prompting Basics
- Zero-shot — ask directly, no examples.
- Few-shot — include 2-5 example input/output pairs in the prompt to demonstrate the desired pattern.
- Chain-of-thought — ask the model to reason step by step before the final answer, often improves accuracy on multi-step problems.
- System prompt — persistent instructions/context set once, applied across the conversation (role, tone, constraints).
- Grounding / RAG (retrieval-augmented generation) — inject retrieved, relevant documents into the prompt so the model answers from supplied facts instead of relying only on parametric memory.
Be specific about format, constraints, and desired length — vague prompts produce vague, inconsistent output.
Vector Embeddings
An embedding maps an input (text, image) to a fixed-length numeric vector such that semantically similar inputs land close together in vector space, measured by cosine similarity or dot product.
# conceptual — using any embedding model
vec_a = embed("a happy dog")
vec_b = embed("a joyful puppy")
vec_c = embed("quarterly tax filing")
# cosine_similarity(vec_a, vec_b) is high — semantically close
# cosine_similarity(vec_a, vec_c) is low — unrelated
Embeddings power semantic search, recommendation, clustering, and retrieval for RAG pipelines — a vector database (e.g. pgvector, Pinecone, Qdrant) stores and indexes them for fast approximate nearest-neighbor lookup at scale.
Tooling Landscape
| Layer | Common tools |
|---|---|
| Numerical / arrays | NumPy |
| Classical ML | scikit-learn |
| Deep learning frameworks | PyTorch, TensorFlow/Keras |
| Data handling | pandas |
| LLM orchestration | LangChain, LlamaIndex |
| Experiment tracking | Weights & Biases, MLflow |
| Model serving | TorchServe, vLLM, Triton, cloud endpoints |
| Notebooks | Jupyter, Google Colab |
For most applied projects: start with scikit-learn for classical ML baselines, reach for PyTorch when you need custom deep learning, and treat LLM API calls (with a lightweight orchestration layer if needed) as the default for most generative-AI use cases before training anything from scratch.