interview

Notes
3 key question in data visualization
Accuracy
Activation Function
Active Learning
Adaboost
AdaBoost vs. Gradient Boosting vs. XGBoost
AdaDelta
AdaGrad
Adam
Adaptive Softmax
Additive Attention
Adjusted R-squared Value
Algorithm - Interview Scheduling
Alternative Hypothesis
Amazon Leadership Principles
Approximate Nearest Neighbor (ANN)
Area Under Curve (AUC)
Area Under Precision Recall Curve (AUPRC)
Auto Regressive Model
Autoencoder for Denoising Images
Averaging in Ensemble Learning
Backward Feature Elimination
Bag of Words
Bagging
Balanced Accuracy
Batch Normalization
Bayes Theorem
Bayesian Optimization Hyperparameter Finding
Beam Search
Behavioral Interview
BERT
BERT Embeddings
Bias & Variance
Bidirectional RNN or LSTM
Binary Cross Entropy
Binning or Bucketing
Binomial Distribution
bitsandbytes
BLEU Score
Boosting
Byte Level BPE
Byte Pair Encoding (BPE)
Causal Language Modeling
Causality vs. Correlation
Central Limit Theorem
Character Tokenizer
Choose the Right Statistical Test
CNN
Co-occurrence based Word Embeddings
Co-Variance
Cold Start Problem
Collaborative Filtering
Collinearity
Conditional Probability
Conditional Random Field
Conditionally Independent Joint Distribution
Confusion Matrix
Contextualized Word Embeddings
Continuous Bag of Words
Continuous Batching
Continuous Random Variable
Contrastive Learning
Contrastive Loss
Convex vs Nonconvex Function
Cosine Similarity
Count based Word Embeddings
Cross Entropy
Cross Validation
Cross-Attention
Crossed Feature
Curse of Dimensionality
Data Augmentation
Data Imputation
Data Visualization
DBScan Clustering
Debugging Deep Learning
Decision Tree
Decision Tree (Classification)
Decision Tree (Regression)
Decoder Only Transformer
Decoding Strategies
Density Sparse Data
Dependent Variable
Derivative
determinant
diagonal-matrix
Differentiation
Differentiation of Product
Digit Dp
Dimensionality Reduction
Discrete Random Variable
Discriminative vs. Generative Models
DistilBERT
Dropout
DS & Algo Interview
Dying ReLU
Dynamic Batching
Dynamic Programming (DP) in python
Elastic Net Regression
Encoder Only Transformer
Encoder-Decoder Transformer
Ensemble Learning
Entropy
Entropy and Information Gain
Essential Visualizations
Estimated Mean
Estimated Standard Deviation
Estimated Variance
Euclidian Norm
Expected Value
Expected Value for Continuous Events
Expected Value for Discrete Events
Exploding Gradient
Exponential Distribution
Extrinsic Evaluation
F-Beta Score
F-Beta@K
F1 Score
False Negative Error
False Positive Rate
FastText Embedding
Feature Engineering
Feature Extraction
Feature Hashing
Feature Preprocessing
Feature Selection
Finding Co-relation between two data or distribution
Forward Feature Selection
Frobenius Norm
Fully Independent Joint Distribution
Fully Joint Distribution
Gaussian Distribution
GBM
Genetic Algorithm Hyperparameter Finding
Gini Impurity
GloVe Embedding
GPU Computation for LLM
Gradient
Gradient Boost (Classification)
Gradient Boost (Regression)
Gradient Boosting
Gradient Checkpointing
Gradient Clipping
Gradient Descent
Graph Convolutional Network (GCN)
Greedy Decoding
Grid Search Hyperparameter Finding
Group Normalization
Group-Query Attention
GRPO
GRU
Gumbel Softmax
Handling Imbalanced Dataset
Handling Missing Data
Handling Outliers
Heapq (nlargest or nsmalles)
Hidden Markov Model
Hierarchical Clustering
Hinge Loss
Histogram
Homonym or Polysemy
How to Choose Kernel in SVM
How to combine in Ensemble Learning
How to prepare for Behavioral Interview
Huber Loss
Hyperparameters
Hypothesis Testing
Identity Matrix
Independent Variable
InfoNCE Loss
Internal Covariate Shift
Interview Resources
Intrinsic Evaluation
Jaccard Distance
Jaccard Similarity
Jacobian Matrix
Joint Distribuition
Joint Probability
K Fold Cross Validation
K-means Clustering
K-means vs. Hierarchical
K-nearest Neighbor (KNN)
Kernel in SVM
Kernel Regression
Kernel Trick
KL Divergence
KTO
KV Cache
L1 or Lasso Regression
L1 vs. L2 Regression
L2 or Ridge Regression
Label Encoding
Label Smoothing
Layer Normalization
Leaky ReLU
Learning Rate Scheduler
Learning to Rank
Leave one out Cross Validation
LightGBM
Likelihood
Linear Regression
Listwise Ranking
Local Attention
Log (Odds Ratio)
Log (Odds)
Log Normalization
Log Scale
Log-cosh Loss
logarithm
Logistic Regression
Logistic Regression vs. Neural Network
LORA
lp-norm
LSTM
Machine Learning Algorithm Selection
Machine Learning vs. Deep Learning
Majority vote in Ensemble Learning
Mamba Architecture
Margin in SVM
Marginal Probability
Masked Self-Attention
Math Dataset
matplotlib legend
Matrices
Max Norm
Maximal Margin Classifier
Maximum Likelihood
Mean
Mean Absolute Error (MAE)
Mean Absolute Percentage Error (MAPE)
Mean Reciprocal Rank (MRR)
Mean Squared Error (MSE)
Mean Squared Logarithmic Error (MSLE)
Median
Merge K-sorted List
Merge Overlapping Intervals
Meteor Score
Min Max Normalization
Mini Batch SGD
Mixed Precision
Mixture of Experts
ML Case Study or ML Design
ML Interview
ML System Design
Mode
Model Based vs. Instance Based Learning
Multi Class Cross Entropy
Multi Label Cross Entropy
Multi Layer Perceptron
Multi-Head Attention
Multi-Head Latent Attention
Multi-Query Attention
Multicollinearity
Multivariable Linear Regression
Multivariate Linear Regression
Multivariate Normal Distribution
Mutual Information
N-gram Method
Naive Bayes
Named Entity Recognition (NER)
NDCG
Negative Log Likelihood
Negative Sampling
Nesterov Accelerated Gradient (NAG)
Neural Network
Neural Network Normalization
Norm of a Vector
Normal Distribution
Null Hypothesis
Odds
Odds Ratio
One Class Classification
One Class Gaussian
One Hot Vector
One vs One Multi Class Classification
One vs Rest or One vs All Multi Class Classification
Optimizers
Optimizing Transformer
Orthogonal Matrix
Orthonormal Vector
Overcomplete Autoencoder
Overfitting
Oversampling
p Value
Padding in CNN
Paged KV Cache
Pairwise Ranking
Parallelism in LLM
Parameter vs. Hyperparameter
PCA vs. Autoencoder
Pearson Correlation
Perceptron
Perplexity
Plots Compared
Pointwise Ranking
Polynomial Regression
Pooling
Population
Positional Encoding in Transformer
Posterior Probability
Pre-Fill in LLM
Pre-Training LLM
Precision
Precision Recall Curve (PRC)
Precision@K
Principal Component Analysis (PCA)
Prior Probability
Probability Density Function
Probability Distribution
Probability Mass Function
Probability vs. Likelihood
Problem Solving Algorithm Selection
Pruning in Decision Tree
PyTorch Refresher
QLORA
Quantization Technique
Questions to ask in a Interview?
Quintile or Percentile
Quotient Rule or Differentiation of Division
R-squared Value
Random Forest
Random Variable
Recall
Recall@K
Recommender System (RecSys)
Regularization
Reinforcement Learning
Reinforcement Learning from Human Feedback (RLHF)
Relational GCN
ReLU
Retrieval Metrics
Retrieval Ranking Funnel
RMSProp
RNN
Robust Scaling Normalization
ROC Curve
Root Mean Squared Error (RMSE)
Root Mean Squared Logarithmic Error (RMSLE)
Rotary Position Embedding (RoPE)
ROUGE-L Score
ROUGE-LSUM Score
ROUGE-N Score
Saddle Points
Self-Attention
Self-Supervised Learning
Semi-supervised Learning
Sensitivity
SentencePiece Tokenization
Sequence-to-Sequence Model
Sigmoid Function
Simple Linear Regression
Singular Value Decomposition (SVD)
Skip Gram Model
Sliding KV Cache
Sliding Window Attention
Soft Margin in SVM
Softmax
Softplus
Softsign
Some Common Behavioral Questions
Sparse Mixture of Experts
Specificity
Splitting tree in Decision Tree
Stacking or Meta Model in Ensemble Learning
Standard deviation
Standardization
Standardization or Normalization
State Space Model
Statistical Significance
Stepwise Selection
Stochastic Gradient Descent (SGD)
Stochastic Gradient Descent with Momentum
Stop Words
Stratified K Fold Cross Validation
Stride in CNN
Stump
Sub-sampling in Word2Vec
Sub-word Tokenizer
Supervised Learning
Support Vector
Support Vector Machine (SVM)
Surprise
SVC
Swallow vs. Deep Learning
Tanh
Temperature in Decoding
Text Preprocessing
TF-IDF
Three Way Partioning
Time Complexity of ML Algos
Time Complexity of ML Models
Tokenizer
Top-K in Retrieval System
Toward RL Learning
Training a Deep Neural Network
Transformer
Transformer vs LSTM
Triplet Loss
True Negative Rate
True Positive Rate
Two-Tower Model
Type 1 Error vs. Type 2 Error
Undercomplete Autoencoder
Undersampling
Unigram Tokenization
Unsupervised Learning
Vanishing Gradient
Vanishing Gradient in Transformers
Variance
Vector Database
Vision Transformer
Weight Initialization
When less data is better than more?
Why do we scale attention weights?
Why do we use Projection in QKV?
Why Trigonometric Function for Positional Encoding?
Word Embeddings
Word Error Rate
Word Tokenizer
Word2Vec Embedding
WordPiece Tokenization
XGBoost
Yet another Rope Extension (YaRN)
Z-Score Normalization