Recent Notes

Note Create At
Binning or Bucketing August 12, 2026
Backward Feature Elimination August 12, 2026
Pooling August 12, 2026
Stride in CNN August 12, 2026
Padding in CNN August 12, 2026
Stump August 12, 2026
Majority vote in Ensemble Learning August 12, 2026
Gradient Boost (Regression) August 12, 2026
Min Max Normalization August 12, 2026
Data Augmentation August 12, 2026
One Hot Vector August 12, 2026
When less data is better than more? August 12, 2026
Density Sparse Data August 12, 2026
Undersampling August 12, 2026
Oversampling August 12, 2026
Log Scale August 12, 2026
Euclidian Norm August 12, 2026
Line Equation August 12, 2026
Chain Rule August 12, 2026
Quotient Rule or Differentiation of Division August 12, 2026
Integration by Parts or Integration of Product August 12, 2026
Differentiation of Product August 12, 2026
Differentiation August 12, 2026
Word Tokenizer August 12, 2026
N-gram Method August 12, 2026
Collinearity August 12, 2026
Character Tokenizer August 12, 2026
RTE (Recognizing Textual Entailment) August 12, 2026
Homonym or Polysemy August 12, 2026
Stop Words August 12, 2026
Unsupervised Learning August 12, 2026
Model Based vs. Instance Based Learning August 12, 2026
Discriminative vs. Generative Models August 12, 2026
Decision Boundary August 12, 2026
Bias & Variance August 12, 2026
Multiset August 12, 2026
Mode August 12, 2026
Median August 12, 2026
Discrete Random Variable August 12, 2026
Continuous Random Variable August 12, 2026
Fully Joint Distribution August 12, 2026
Conditionally Independent Joint Distribution August 12, 2026
Quintile or Percentile August 12, 2026
Co-Variance August 12, 2026
Estimated Variance August 12, 2026
Estimated Mean August 12, 2026
Exponential Distribution August 12, 2026
Expected Value for Discrete Events August 12, 2026
Entropy and Information Gain August 12, 2026
Surprise August 12, 2026
Odds Ratio August 12, 2026
Odds August 12, 2026
Likelihood August 12, 2026
Probability vs. Likelihood August 12, 2026
Prior Probability August 12, 2026
Marginal Probability August 12, 2026
Conditional Probability August 12, 2026
Why do we use Projection in QKV? August 12, 2026
Toward RL Learning August 12, 2026
Reinforcement Learning August 12, 2026
Paged KV Cache August 12, 2026
Untitled August 12, 2026
Soft Margin in SVM August 11, 2026
What is More Likely to Happen Next August 11, 2026
tau-bench - A Benchmark for Tool-Agent-User Interaction in Real-World Domains August 11, 2026
DeepSeek-R1 August 11, 2026
Molmo and PixMo August 11, 2026
Deliberative Alignment - Reasoning Enables Safer Language Models August 11, 2026
Token Assorted - Mixing Latent and Text Tokens for Improved Language Model Reasoning August 11, 2026
G-Eval - NLG Evaluation using GPT-4 with Better Human Alignment August 11, 2026
Vision Language Model-based Caption Evaluation Method Leveraging Visual Context Extraction August 11, 2026
Large Language Models are Zero-Shot Rankers for Recommender Systems August 11, 2026
Investigating Continual Pretraining in Large Language Models - Insights and Implications August 11, 2026
Piecing It All Together - Verifying Multi-Hop Multimodal Claims August 11, 2026
Is a Question Decomposition Unit All We Need August 11, 2026
Compressed Chain of Thought - Efficient Reasoning Through Dense Representations August 11, 2026
OpenPI-C August 11, 2026
Scientific Fact-Checking - A Survey of Resources and Approaches August 11, 2026
Semantic Product Search for Matching Structured Product Catalogs in E-Commerce August 11, 2026
PubMedQA - A Dataset for Biomedical Research Question Answering August 11, 2026
Multi-Head Latent Attention August 11, 2026
Group Normalization August 11, 2026
Jaccard Similarity August 11, 2026
Decision Tree August 11, 2026
logarithm August 11, 2026
Instructional Websites August 11, 2026
DistilBERT August 11, 2026
Named Entity Recognition (NER) August 11, 2026
Type 1 Error vs. Type 2 Error August 11, 2026
Polynomial Regression August 11, 2026
Encoder-Decoder Transformer August 11, 2026
Byte Level BPE August 11, 2026
Continuous Bag of Words August 11, 2026
Dynamic Programming (DP) in python August 11, 2026
How to do research August 11, 2026
Math Dataset August 11, 2026
GPT-OSS August 11, 2026
Precision Recall Curve (PRC) August 11, 2026
Reno Talk @UMBC on Scale-2024 August 11, 2026
Independent Variable August 11, 2026
matplotlib functions August 11, 2026
jupyter-notebook-on-server August 11, 2026
Gradient Clipping August 11, 2026
ROUGE-LSUM Score August 11, 2026
BLEU Score August 11, 2026
F-Beta Score August 11, 2026
RMSProp August 11, 2026
Time Complexity of ML Algos August 11, 2026
Crossed Feature August 11, 2026
Claude Code August 11, 2026
Multi-Query Attention August 11, 2026
Time Complexity of ML Models August 11, 2026
Sliding Window Attention August 11, 2026
Pre-Fill in LLM August 11, 2026
FastText Embedding August 11, 2026
Skip Gram Model August 11, 2026
Extrinsic Evaluation August 11, 2026
Feature Hashing August 11, 2026
Area Under Precision Recall Curve (AUPRC) August 11, 2026
GloVe Embedding August 11, 2026
Jaccard Distance August 11, 2026
Forward Feature Selection August 11, 2026
do_sample (vllm vs hf) August 11, 2026
Why do we scale attention weights? August 11, 2026
Balanced Accuracy August 11, 2026
Claim Verification Datasets August 11, 2026
Group-Query Attention August 11, 2026
Support Vector Machine (SVM) August 11, 2026
Zotero Guides August 11, 2026
F-Beta@K August 11, 2026
Adjusted R-squared Value August 11, 2026
Activation Function August 11, 2026
Fake News Challenge August 11, 2026
Research Poster August 11, 2026
Count based Word Embeddings August 11, 2026
ELMo Embeddings August 11, 2026
Negative Sampling August 11, 2026
Mean Reciprocal Rank (MRR) August 11, 2026
Fine Tuning Large Language Models August 11, 2026
Leaky ReLU August 11, 2026
Sub-sampling in Word2Vec August 11, 2026
Prepare for Talk August 11, 2026
How to prepare for Behavioral Interview August 11, 2026
Intrinsic Evaluation August 11, 2026
Alternative Hypothesis August 11, 2026
Word Embeddings August 11, 2026
Linear Regression with Normal Equation August 11, 2026
Research Skills Unsorted List August 11, 2026
matplotlib legend August 11, 2026
Retrieval Metrics August 11, 2026
Self-Supervised Learning August 11, 2026
Research Slides or Research Talk August 11, 2026
Neural Network Normalization August 11, 2026
Causal Language Modeling August 11, 2026
Label Encoding August 11, 2026
Basics of Kubernetes August 11, 2026
Top-K in Retrieval System August 11, 2026
Personal Claude Code Recommendations August 11, 2026
Multivariate Linear Regression August 11, 2026
True Negative Rate August 11, 2026
Word2Vec Embedding August 11, 2026
Presentation Making Tips August 11, 2026
Dependent Variable August 11, 2026
Feature Preprocessing August 11, 2026
Curse of Dimensionality August 11, 2026
ML Case Study or ML Design August 11, 2026
LLM GPU Calculate August 11, 2026
Simple Linear Regression August 11, 2026
Multivariable Linear Regression August 11, 2026
Local Attention August 11, 2026
Sequence-to-Sequence Model August 11, 2026
Tokenizer August 11, 2026
Foundation Model August 11, 2026
3 key question in data visualization August 11, 2026
Co-occurrence based Word Embeddings August 11, 2026
Log (Odds Ratio) August 11, 2026
doing-literature-review August 11, 2026
MM-LLMs August 11, 2026
spacy-semantic-similarity August 11, 2026
spacy-pos August 11, 2026
spacy-pipeline August 11, 2026
spacy-pattern August 11, 2026
spacy-operator-quantifier August 11, 2026
spacy-named-entities August 11, 2026
spacy-matcher August 11, 2026
spacy-explanation-of-labels August 11, 2026
spacy-doc-span-token August 11, 2026
spacy-doc-object August 11, 2026
Joint Distribuition August 11, 2026
Scalar August 11, 2026
Orthonormal Vector August 11, 2026
Orthogonal Matrix August 11, 2026
Norm of a Vector August 11, 2026
Max Norm August 11, 2026
Matrices August 11, 2026
lp-norm August 11, 2026
Identity Matrix August 11, 2026
Frobenius Norm August 11, 2026
diagonal-matrix August 11, 2026
determinant August 11, 2026
TF-IDF August 11, 2026
Sensitivity August 11, 2026
Recall August 11, 2026
Precision August 11, 2026
F1 Score August 11, 2026
Accuracy August 11, 2026
COIN August 11, 2026
Swallow vs. Deep Learning August 11, 2026
ROUGE-N Score August 11, 2026
Parameter vs. Hyperparameter August 11, 2026
Normal Distribution August 11, 2026
Naive Bayes August 11, 2026
Multivariate Normal Distribution August 11, 2026
Multi Layer Perceptron August 11, 2026
Meteor Score August 11, 2026
Machine Learning Algorithm Selection August 11, 2026
LSTM August 11, 2026
Log (Odds) August 11, 2026
L1 or Lasso Regression August 11, 2026
KL Divergence August 11, 2026
Gradient Boosting August 11, 2026
Gini Impurity August 11, 2026
Genetic Algorithm Hyperparameter Finding August 11, 2026
Exploding Gradient August 11, 2026
Entropy August 11, 2026
Cross Validation August 11, 2026
Convex vs Nonconvex Function August 11, 2026
CNN August 11, 2026
BERT August 11, 2026
Bagging August 11, 2026
Area Under Curve (AUC) August 11, 2026
Two-Tower Model August 10, 2026
Recommender System (RecSys) August 10, 2026
Retrieval Ranking Funnel August 10, 2026
NDCG August 10, 2026
Pairwise Ranking August 10, 2026
Pointwise Ranking August 10, 2026
Learning to Rank August 10, 2026
Listwise Ranking August 10, 2026
Collaborative Filtering August 10, 2026
Cold Start Problem August 10, 2026
Approximate Nearest Neighbor (ANN) August 10, 2026
spacy-syntactic-dependency August 10, 2026
p Value August 10, 2026
bisect_left vs. bisect_right August 10, 2026
WordPiece Tokenization August 10, 2026
Why Trigonometric Function for Positional Encoding? August 10, 2026
Variance August 10, 2026
Undercomplete Autoencoder August 10, 2026
Training a Deep Neural Network August 10, 2026
Transformer vs LSTM August 10, 2026
Triplet Loss August 10, 2026
True Positive Rate August 10, 2026
Three Way Partioning August 10, 2026
Tanh August 10, 2026
Sub-word Tokenizer August 10, 2026
Supervised Learning August 10, 2026
Support Vector August 10, 2026
Stochastic Gradient Descent (SGD) August 10, 2026
Stochastic Gradient Descent with Momentum August 10, 2026
Standardization August 10, 2026
Statistical Significance August 10, 2026
Specificity August 10, 2026
Splitting tree in Decision Tree August 10, 2026
Stacking or Meta Model in Ensemble Learning August 10, 2026
Standard deviation August 10, 2026
Softmax August 10, 2026
Softplus August 10, 2026
Softsign August 10, 2026
Some Common Behavioral Questions August 10, 2026
Sources of Uncertainty August 10, 2026
SentencePiece Tokenization August 10, 2026
Self-Attention August 10, 2026
Semi-supervised Learning August 10, 2026
STORM Method August 10, 2026
SVC August 10, 2026
Saddle Points August 10, 2026
Second Order Derivative or Hessian Matrix August 10, 2026
Root Mean Squared Error (RMSE) August 10, 2026
Root Mean Squared Logarithmic Error (RMSLE) August 10, 2026
Regularization August 10, 2026
Reinforcement Learning from Human Feedback (RLHF) August 10, 2026
Random Variable August 10, 2026
ReLU August 10, 2026
Recall@K August 10, 2026
Random Forest August 10, 2026
R-squared Value August 10, 2026
RNN August 10, 2026
ROC Curve August 10, 2026
ROUGE-L Score August 10, 2026
Questions to ask in a Interview? August 10, 2026
PyTorch Refresher August 10, 2026
Problem Solving Algorithm Selection August 10, 2026
PyTorch Loss Functions August 10, 2026
Principal Component Analysis (PCA) August 10, 2026
Probability Density Function August 10, 2026
Probability Distribution August 10, 2026
Probability Mass Function August 10, 2026
Pre-Training LLM August 10, 2026
Precision@K August 10, 2026
Population August 10, 2026
Positional Encoding in Transformer August 10, 2026
Posterior Probability August 10, 2026
Perceptron August 10, 2026
Perplexity August 10, 2026
PCA vs. Autoencoder August 10, 2026
Optimizing Transformer August 10, 2026
Overcomplete Autoencoder August 10, 2026
One vs One Multi Class Classification August 10, 2026
One vs Rest or One vs All Multi Class Classification August 10, 2026
One Class Classification August 10, 2026
One Class Gaussian August 10, 2026
Null Hypothesis August 10, 2026
Normalization August 10, 2026
Neural Network August 10, 2026
Negative Log Likelihood August 10, 2026
Mutual Information August 10, 2026
Multi Label Cross Entropy August 10, 2026
Multi-Head Attention August 10, 2026
Multi Class Cross Entropy August 10, 2026
Mini Batch SGD August 10, 2026
Mean August 10, 2026
Merge K-sorted List August 10, 2026
Merge Overlapping Intervals August 10, 2026
Mean Absolute Error (MAE) August 10, 2026
Mean Absolute Percentage Error (MAPE) August 10, 2026
Mean Squared Error (MSE) August 10, 2026
Mean Squared Logarithmic Error (MSLE) August 10, 2026
Maximal Margin Classifier August 10, 2026
Maximum Likelihood August 10, 2026
Masked Self-Attention August 10, 2026
Margin in SVM August 10, 2026
Machine Learning vs. Deep Learning August 10, 2026
Log-cosh Loss August 10, 2026
Logistic Regression vs. Neural Network August 10, 2026
Local Minima August 10, 2026
Layer Normalization August 10, 2026
Kernel in SVM August 10, 2026
L2 or Ridge Regression August 10, 2026
L1 vs. L2 Regression August 10, 2026
Kernel Regression August 10, 2026
KV Cache August 10, 2026
K Fold Cross Validation August 10, 2026
K-means Clustering August 10, 2026
K-means vs. Hierarchical August 10, 2026
K-nearest Neighbor (KNN) August 10, 2026
Interview Resources August 10, 2026
Instruction Fine Tuning August 10, 2026
Hyperparameters August 10, 2026
Huber Loss August 10, 2026
How to Choose Kernel in SVM August 10, 2026
How to Write Academic Paper (from CS Perspective) August 10, 2026
How to combine in Ensemble Learning August 10, 2026
Hierarchical Clustering August 10, 2026
Histogram August 10, 2026
Heapq (nlargest or nsmalles) August 10, 2026
Grid Search Hyperparameter Finding August 10, 2026
Gumbel Softmax August 10, 2026
Handling Imbalanced Dataset August 10, 2026
Gradient August 10, 2026
Gradient Descent August 10, 2026
Greedy Decoding August 10, 2026
Gradient Boost (Classification) August 10, 2026
Global Minima August 10, 2026
GPU Computation for LLM August 10, 2026
GRU August 10, 2026
Gaussian Distribution August 10, 2026
GBM August 10, 2026
Finding Co-relation between two data or distribution August 10, 2026
False Positive Rate August 10, 2026
Feature Engineering August 10, 2026
Feature Extraction August 10, 2026
False Negative Error August 10, 2026
Expected Value August 10, 2026
Estimated Standard Deviation August 10, 2026
Encoder Only Transformer August 10, 2026
Ensemble Learning August 10, 2026
Essential Visualizations August 10, 2026
Eigendecomposition August 10, 2026
Elastic Net Regression August 10, 2026
Domain vs. Codomain vs. Range August 10, 2026
Dropout August 10, 2026
Dying ReLU August 10, 2026
Decoder Only Transformer August 10, 2026
Decision Tree (Regression) August 10, 2026
Derivative August 10, 2026
Data Imputation August 10, 2026
Decision Tree (Classification) August 10, 2026
Cross-Attention August 10, 2026
DBScan Clustering August 10, 2026
Contrastive Learning August 10, 2026
Contrastive Loss August 10, 2026
Cosine Similarity August 10, 2026
Cross Entropy August 10, 2026
Contextualized Word Embeddings August 10, 2026
Confusion Matrix August 10, 2026
Challenges of NLP (2022) August 10, 2026
Causality vs. Correlation August 10, 2026
Byte Pair Encoding (BPE) August 10, 2026
Binary Cross Entropy August 10, 2026
Binomial Distribution August 10, 2026
Boosting August 10, 2026
Bidirectional RNN or LSTM August 10, 2026
Batch Normalization August 10, 2026
Bayes Theorem August 10, 2026
Beam Search August 10, 2026
Behavioral Interview August 10, 2026
Bag of Words August 10, 2026
Averaging in Ensemble Learning August 10, 2026
BERT Embeddings August 10, 2026
Auto Regressive Model August 10, 2026
Autoencoder for Denoising Images August 10, 2026
Additive Attention August 10, 2026
Amazon Leadership Principles August 10, 2026
AdaGrad August 10, 2026
Adaboost August 10, 2026
Adam August 10, 2026
AdaDelta August 10, 2026
Recent Notes August 10, 2026
MultiVENT August 10, 2026
How To Write a Paper August 10, 2026
How to Read a Paper August 10, 2026
Advanced NLP with Scipy August 10, 2026
Deep Learning by Ian Goodfellow August 10, 2026
How To 100M Learning Text Video August 10, 2026
Home August 10, 2026