Two-Tower Model
- The workhorse of modern candidate generation (stage 1 of the Retrieval Ranking Funnel).
- Two separate nets ("towers"):
- User/query tower → one vector
u(from user history, context, or query text). - Item tower → one vector
v(from item features), in the same space.
- User/query tower → one vector
- Score of a (user, item) pair = dot product
u · v(or cosine).
Why it scales (the whole point)
- Item vector
vdoes not depend on the user. - So precompute
vfor all items offline, store in a vector index. - Online: one user forward pass for
u, then a nearest-neighbor lookup. Billions of items, but only one forward pass + a lookup per request. - Exact nearest-neighbor over billions is still too slow, so use Approximate Nearest Neighbor (ANN).
The constraint
- Two-tower can only use features that factor into a pure-user side and a pure-item side.
- Cross features (user × item, e.g. "times this user clicked this item's category in the last hour") depend on the pair, so they break the user-independence of
v. You'd have to score every pair online, which kills the speed. - Cross features therefore belong in the ranking stage (Learning to Rank), where the candidate set is already small.
Relation to matrix factorization
- Same shape as Collaborative Filtering via Matrix Factorization, but with learned neural towers instead of plain embedding lookups. Neural towers can take features, so they handle new items (Cold Start Problem).