ML And AI Basics

Core Terminology

TermMeaning
FeatureAn input variable used to make a prediction
Label / targetThe value a supervised model learns to predict
ModelA function (with learned parameters) mapping inputs to outputs
ParametersValues learned during training (e.g. weights)
HyperparametersValues set before training (learning rate, layer count) — not learned
InferenceRunning a trained model on new input to get a prediction
Dataset splitTraining (fit), validation (tune), test (final unbiased check)

Types Of Learning

  • Supervised — learn a mapping from labeled input→output pairs (classification, regression).
  • Unsupervised — find structure in unlabeled data (clustering, dimensionality reduction).
  • Self-supervised — generate labels from the data itself (e.g. predict the next token) — how most large language models are pretrained.
  • Reinforcement learning — an agent learns by taking actions in an environment and receiving reward signals (used in game-playing agents, RLHF fine-tuning).

Model Training Loop

for epoch in range(num_epochs):
    for batch in train_loader:
        predictions = model(batch.inputs)
        loss = loss_fn(predictions, batch.labels)
        loss.backward()          # compute gradients (backpropagation)
        optimizer.step()          # update weights
        optimizer.zero_grad()      # reset gradients for next batch
    validate(model, val_loader)      # check generalization each epoch

Backpropagation computes how much each parameter contributed to the error via the chain rule; the optimizer (e.g. SGD, Adam) nudges parameters to reduce loss. An epoch is one full pass over the training data; a batch is a subset processed per gradient step.

Evaluation Metrics

TaskMetrics
Classificationaccuracy, precision, recall, F1, ROC-AUC, confusion matrix
RegressionMAE, MSE, RMSE, R²
Ranking / retrievalprecision@k, recall@k, NDCG, MRR
Generation (LLM)perplexity, BLEU/ROUGE (translation/summarization), human/LLM-judged eval

Precision = of predicted positives, how many were actually positive. Recall = of actual positives, how many were caught. F1 is their harmonic mean — use it when classes are imbalanced and accuracy alone is misleading.

Overfitting And Regularization

Overfitting: the model memorizes training data (including noise) instead of learning generalizable patterns — low training loss, high validation loss.

Mitigations:

  • More/varied data — the most reliable fix.
  • Regularization — L1/L2 weight penalties, dropout (randomly zero activations during training).
  • Early stopping — stop training when validation loss stops improving.
  • Simpler model / fewer parameters — reduce capacity to match the data available.
  • Cross-validation — check generalization isn’t a fluke of one train/val split.

Underfitting is the opposite failure — the model is too simple to capture the pattern (high error on both train and validation).

Neural Network Basics

A neural network is layers of linear transformations (weights + bias) interleaved with nonlinear activation functions, trained end-to-end via gradient descent.

input → [linear → activation] × N layers → output
ActivationTypical use
ReLUhidden layers (default choice, cheap, avoids vanishing gradients)
Sigmoidbinary classification output (0-1 probability)
Softmaxmulti-class classification output (probability distribution)
Tanhoccasionally in RNNs/older architectures

Without nonlinear activations, stacking linear layers collapses to a single linear function — depth would add no expressive power.

Common Architectures

ArchitectureGood for
CNN (convolutional)images, spatial/grid data — learns local patterns via shared filters
RNN / LSTM / GRUsequences (legacy for NLP/time series) — processes step by step, struggles with long-range dependencies
Transformersequences, now dominant for NLP and increasingly vision — attention lets every position see every other position directly
Autoencodercompression, anomaly detection, denoising — learns to reconstruct its input through a narrow bottleneck
GAN / diffusiongenerative image/audio synthesis

LLM Basics

Large language models are transformer-based, trained to predict the next token over massive text corpora (self-supervised pretraining), then adjusted via instruction tuning and RLHF/DPO to follow instructions and align with human preference.

  • Token — a chunk of text (roughly a word-piece or subword) the model actually processes; not the same as a character or whole word.
  • Context window — the maximum number of tokens a model can attend to in one request (input + output combined).
  • Temperature — sampling randomness at inference: near 0 is deterministic/focused, higher values are more varied/creative.
  • Hallucination — the model generates plausible-sounding but factually wrong or unsupported output — an inherent risk since generation is a probability model, not a lookup.

Prompting Basics

  • Zero-shot — ask directly, no examples.
  • Few-shot — include 2-5 example input/output pairs in the prompt to demonstrate the desired pattern.
  • Chain-of-thought — ask the model to reason step by step before the final answer, often improves accuracy on multi-step problems.
  • System prompt — persistent instructions/context set once, applied across the conversation (role, tone, constraints).
  • Grounding / RAG (retrieval-augmented generation) — inject retrieved, relevant documents into the prompt so the model answers from supplied facts instead of relying only on parametric memory.

Be specific about format, constraints, and desired length — vague prompts produce vague, inconsistent output.

Vector Embeddings

An embedding maps an input (text, image) to a fixed-length numeric vector such that semantically similar inputs land close together in vector space, measured by cosine similarity or dot product.

# conceptual — using any embedding model
vec_a = embed("a happy dog")
vec_b = embed("a joyful puppy")
vec_c = embed("quarterly tax filing")
# cosine_similarity(vec_a, vec_b) is high — semantically close
# cosine_similarity(vec_a, vec_c) is low — unrelated

Embeddings power semantic search, recommendation, clustering, and retrieval for RAG pipelines — a vector database (e.g. pgvector, Pinecone, Qdrant) stores and indexes them for fast approximate nearest-neighbor lookup at scale.

Tooling Landscape

LayerCommon tools
Numerical / arraysNumPy
Classical MLscikit-learn
Deep learning frameworksPyTorch, TensorFlow/Keras
Data handlingpandas
LLM orchestrationLangChain, LlamaIndex
Experiment trackingWeights & Biases, MLflow
Model servingTorchServe, vLLM, Triton, cloud endpoints
NotebooksJupyter, Google Colab

For most applied projects: start with scikit-learn for classical ML baselines, reach for PyTorch when you need custom deep learning, and treat LLM API calls (with a lightweight orchestration layer if needed) as the default for most generative-AI use cases before training anything from scratch.