You have a working chat() call and you want it to remember context across turns or sessions. By the end of this guide, memoryMiddleware recalls relevant memory into the prompt and saves each finished turn through a real adapter, scoped safely from your server-validated session.
Want the full contract first? See the Overview.
pnpm add @tanstack/ai-memorypnpm add @tanstack/ai-memory@tanstack/ai-memory ships memoryMiddleware, the MemoryAdapter contract, and the built-in and vendor adapters (each on its own subpath).
In-memory: inMemory() is zero-dependency and stores records in a Map. Use it for local development, tests, and single-process demos. Records vanish on restart.
Redis: redis({ redis }) persists across restarts and shares state across processes. Bring your own client (ioredis, or redis via fromNodeRedis).
Vendors: hindsight(), mem0(), and honcho() delegate to a hosted memory service.
Custom adapters implement the recall/save contract. See Custom Adapter.
Start with the in-memory adapter, the fastest path to a working setup:
import { chat } from '@tanstack/ai'
import { openaiText } from '@tanstack/ai-openai'
import { memoryMiddleware } from '@tanstack/ai-memory'
import { inMemory } from '@tanstack/ai-memory/in-memory'
const memory = inMemory()
const stream = chat({
adapter: openaiText('gpt-5.5'),
messages: [{ role: 'user', content: 'Hello' }],
middleware: [
memoryMiddleware({
adapter: memory,
scope: { sessionId: 'demo-thread', userId: 'alice' },
}),
],
})import { chat } from '@tanstack/ai'
import { openaiText } from '@tanstack/ai-openai'
import { memoryMiddleware } from '@tanstack/ai-memory'
import { inMemory } from '@tanstack/ai-memory/in-memory'
const memory = inMemory()
const stream = chat({
adapter: openaiText('gpt-5.5'),
messages: [{ role: 'user', content: 'Hello' }],
middleware: [
memoryMiddleware({
adapter: memory,
scope: { sessionId: 'demo-thread', userId: 'alice' },
}),
],
})Each turn, the middleware recalls relevant memory into the system prompt (lexical scoring by default), then deferred-saves the user and assistant turn after the stream finishes.
When you're ready to ship, swap the adapter and keep everything else the same:
import Redis from 'ioredis'
import { memoryMiddleware } from '@tanstack/ai-memory'
import { redis } from '@tanstack/ai-memory/redis'
import type { MemoryScope } from '@tanstack/ai-memory'
declare const scope: MemoryScope // from Step 5
const client = new Redis(process.env.REDIS_URL ?? 'redis://localhost:6379')
const memory = redis({ redis: client })
memoryMiddleware({ adapter: memory, scope })import Redis from 'ioredis'
import { memoryMiddleware } from '@tanstack/ai-memory'
import { redis } from '@tanstack/ai-memory/redis'
import type { MemoryScope } from '@tanstack/ai-memory'
declare const scope: MemoryScope // from Step 5
const client = new Redis(process.env.REDIS_URL ?? 'redis://localhost:6379')
const memory = redis({ redis: client })
memoryMiddleware({ adapter: memory, scope })Using a hosted service? Swap inMemory() for hindsight({ user }), mem0({ user }), or honcho({ user }). The middleware wiring is identical. The adapter maps recall/save onto the vendor API.
The built-in adapters score lexically by default. Pass an embedder for semantic recall when scopes grow large or queries don't share keywords with stored text:
import OpenAI from 'openai'
import { inMemory } from '@tanstack/ai-memory/in-memory'
const openai = new OpenAI()
const memory = inMemory({
embedder: {
async embed(text) {
const result = await openai.embeddings.create({
model: 'text-embedding-3-small',
input: text,
})
const embedding = result.data[0]?.embedding
if (!embedding) throw new Error('embedding request returned no vector')
return embedding
},
},
})import OpenAI from 'openai'
import { inMemory } from '@tanstack/ai-memory/in-memory'
const openai = new OpenAI()
const memory = inMemory({
embedder: {
async embed(text) {
const result = await openai.embeddings.create({
model: 'text-embedding-3-small',
input: text,
})
const embedding = result.data[0]?.embedding
if (!embedding) throw new Error('embedding request returned no vector')
return embedding
},
},
})scope is the isolation boundary. Static scopes are fine for fixtures, but in any real app derive scope per request from server-validated session data, never from the request body.
import { chat } from '@tanstack/ai'
import { openaiText } from '@tanstack/ai-openai'
import { memoryMiddleware } from '@tanstack/ai-memory'
import type { ModelMessage } from '@tanstack/ai'
import type { MemoryAdapter } from '@tanstack/ai-memory'
// From earlier steps / your auth layer.
declare const messages: Array<ModelMessage>
declare const memory: MemoryAdapter
declare const session: { userId: string; threadId: string }
declare function getSession(ctx: unknown): { threadId: string; userId: string }
const stream = chat({
adapter: openaiText('gpt-5.5'),
messages,
context: { session }, // attached by your auth middleware, not from req.body
middleware: [
memoryMiddleware({
adapter: memory,
scope: (ctx) => {
const session = getSession(ctx)
return { sessionId: session.threadId, userId: session.userId }
},
}),
],
})import { chat } from '@tanstack/ai'
import { openaiText } from '@tanstack/ai-openai'
import { memoryMiddleware } from '@tanstack/ai-memory'
import type { ModelMessage } from '@tanstack/ai'
import type { MemoryAdapter } from '@tanstack/ai-memory'
// From earlier steps / your auth layer.
declare const messages: Array<ModelMessage>
declare const memory: MemoryAdapter
declare const session: { userId: string; threadId: string }
declare function getSession(ctx: unknown): { threadId: string; userId: string }
const stream = chat({
adapter: openaiText('gpt-5.5'),
messages,
context: { session }, // attached by your auth middleware, not from req.body
middleware: [
memoryMiddleware({
adapter: memory,
scope: (ctx) => {
const session = getSession(ctx)
return { sessionId: session.threadId, userId: session.userId }
},
}),
],
})On the client, nothing changes. useChat (or your connection adapter) consumes the stream exactly as before. Memory is entirely server-side.