A MAP hypervector bundle combines bound facts. This repository tests that bundle from two directions.
Throughout these docs,
| Inside the bundle | Outside the bundle | |
|---|---|---|
| Question | How many facts can one bundle hold before we can't read a fact back? | How many records fit before chance alone makes two unrelated records look like a match, for one search and for the whole dataset? |
| Pressure | Crowding within one vector: every added fact is noise for the others | Crowding among vectors: every added candidate is another chance to score high by accident |
| Workload | One person's scored answers to 5–50 IPIP questionnaire items | 40,000 synthetic records with five categorical properties: region, education, occupation, interest cluster and employer |
| Encoder |
|
|
| Measures | Answer cleanup, whole-profile geometry, similarity as bundles grow | False matches and missed matches against a match threshold, for one search and for the whole dataset, at D = 2,048, 4,096 and 8,192 |
| Report | src/qa-encoding/REPORT.md | src/scale/REPORT.md |
| Code | src/qa-encoding |
src/scale · methodology
|
| Data |
data/qa-encoding, 31 fixed synthetic profiles |
data/scale, one fixed sequence of uniformly drawn records |
The two studies use different encoders and workloads, so their numbers should not be compared as if one encoder produced both. Both datasets are synthetic. The original records stay the exact source of truth; bundle readout and similarity are measured properties of each encoder on its data, not general capacity limits of a hyperspace.
Each study has its own locked environment under src/<study>/. See each README for the reproduction commands.