03 / ALOA · Machine learning security
What does amodel remember?
Agnostic membership inference on two-tower neural networks. The question is whether a record belonged to training—not how to reconstruct it.
Separate user and item towers produce embeddings.
Follow the full method.
Separate user and item inputs become embeddings.
Synthetic queries train a shadow model on target responses.
User-feature perturbations expose response behavior.
Shadow-derived features support membership classification.
Read the result precisely.
Table 6 reports 97.87% Combined accuracy and 95.07% Dummy accuracy with IN:OUT ratios of 4:1 and 1:10. These populations differ. Table 7’s 99.65% is the proportion of known training members predicted IN, not general member/nonmember accuracy.
Open the full paper and discussionWhat the diagrams do—and do not—show
The positions, pulses, and embedding shifts are conceptual. They are not measured coordinates or an empirical ROC curve. Membership inference asks about training inclusion; it does not demonstrate record reconstruction. MMD distribution comparisons should not be read as membership probabilities.