Response Compaction
7/28/26About 2 min
Response Compaction
Response Compaction compresses long conversation histories into concise summaries, significantly reducing token usage for subsequent requests. Ideal for long-running dialogues that exceed context window limits.
Endpoint Info
| Item | Value |
|---|---|
| URL | /v1/responses/compact |
| Method | POST |
| Auth | Bearer Token (API Key) |
| Content-Type | application/json |
Request Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
model | string | Yes | Model ID, e.g. gpt-4o-mini |
input | array | Yes | Message array (full conversation history) |
max_output_tokens | integer | No | Maximum tokens for the compressed output |
instructions | string | No | Custom instructions for how to compress |
input format
The input field accepts a message array in standard chat format:
{
"input": [
{"role": "user", "content": "What is machine learning?"},
{"role": "assistant", "content": "Machine learning is a branch of AI..."},
{"role": "user", "content": "How does it differ from deep learning?"},
{"role": "assistant", "content": "Deep learning is a subset of machine learning..."}
]
}cURL Examples
Basic Compaction
curl https://api.quickapi.store/v1/responses/compact \
-H "Content-Type: application/json" \
-H "Authorization: Bearer YOUR_API_KEY" \
-d '{
"model": "gpt-4o-mini",
"input": [
{"role": "user", "content": "What is machine learning?"},
{"role": "assistant", "content": "Machine learning is a branch of artificial intelligence that enables systems to learn from data."},
{"role": "user", "content": "How does it differ from deep learning?"},
{"role": "assistant", "content": "Deep learning is a subset of machine learning that uses neural networks with many layers."},
{"role": "user", "content": "What are real-world applications?"},
{"role": "assistant", "content": "Applications include image recognition, natural language processing, recommendation systems, and autonomous driving."}
],
"max_output_tokens": 500
}'With Custom Instructions
curl https://api.quickapi.store/v1/responses/compact \
-H "Content-Type: application/json" \
-H "Authorization: Bearer YOUR_API_KEY" \
-d '{
"model": "gpt-4o",
"input": [
{"role": "user", "content": "Explain the project architecture"},
{"role": "assistant", "content": "The project uses a microservices architecture with three main layers..."},
{"role": "user", "content": "What about the database layer?"},
{"role": "assistant", "content": "We use PostgreSQL for relational data and Redis for caching..."}
],
"instructions": "Preserve all technical decisions and architecture details. Omit pleasantries.",
"max_output_tokens": 300
}'💡 When to Use Compaction
Use compaction when your conversation exceeds 50% of the model's context window. The compressed summary preserves key information while reducing token count by 60-80%.
Response Example
{
"id": "compact_abc123",
"object": "response.compaction",
"created_at": 1700000000,
"model": "gpt-4o-mini",
"output": {
"type": "message",
"role": "assistant",
"content": [
{
"type": "output_text",
"text": "Conversation Summary:\n\nThe user asked about machine learning (ML), deep learning (DL), and their applications. Key points covered:\n- ML is an AI branch that learns from data\n- DL is a subset of ML using multi-layer neural networks\n- Real-world applications: image recognition, NLP, recommendations, autonomous driving"
}
]
},
"usage": {
"input_tokens": 850,
"output_tokens": 72,
"compression_ratio": 0.91,
"total_tokens": 922
}
}Python Example
from openai import OpenAI
client = OpenAI(
base_url="https://api.quickapi.store/v1",
api_key="YOUR_API_KEY"
)
# Compact a long conversation
messages = [
{"role": "user", "content": "What is machine learning?"},
{"role": "assistant", "content": "Machine learning is a branch of artificial intelligence that enables systems to learn from data."},
{"role": "user", "content": "How does it differ from deep learning?"},
{"role": "assistant", "content": "Deep learning is a subset of machine learning that uses neural networks with many layers."},
{"role": "user", "content": "What are real-world applications?"},
{"role": "assistant", "content": "Applications include image recognition, natural language processing, recommendation systems, and autonomous driving."}
]
response = client.responses.create(
model="gpt-4o-mini",
input=messages,
instructions="Compress this conversation into a concise summary"
)
# Use the compacted summary as context for future turns
compressed = response.output[0].content[0].text
print(f"Compressed: {compressed}")
print(f"Compression ratio: {response.usage.compression_ratio}")⚠️ Information Loss
Compression inherently discards some detail. For critical conversations where every detail matters, consider using a model with a larger context window instead of compaction.

