Hermes memory benchmark on MemConflict
A comparison of self-hostable memory providers for the Hermes agent, measured on the MemConflict benchmark: 30 personas and 3,750 questions about changing, contradictory, and condition-bound memories.
Report
Benchmark results
Per-provider accuracy, retrieval quality, failure analysis, workload cost, and the method behind the numbers.
DatasetConversation browser
Read the MemConflict dialogues as chat transcripts, with each session's conflict annotations and questions.
The conversation browser loads the full dataset as one 40 MB script. Give it a moment on a slow connection.
Source documents
- BENCHMARK_MATRIX.md: providers, arms, flags, and every measured number.
- DECISIONS.md: why the harness is built this way, including reversed decisions.
- TROUBLESHOOTING.md: symptom, cause, fix, and what did not work.