M/7

Hermes memory benchmark on MemConflict

A comparison of self-hostable memory providers for the Hermes agent, measured on the MemConflict benchmark: 30 personas and 3,750 questions about changing, contradictory, and condition-bound memories.

Report

Benchmark results

Per-provider accuracy, retrieval quality, failure analysis, workload cost, and the method behind the numbers.

Dataset

Conversation browser

Read the MemConflict dialogues as chat transcripts, with each session's conflict annotations and questions.

The conversation browser loads the full dataset as one 40 MB script. Give it a moment on a slow connection.

Source documents