Real-world observation: OpenAI embedding API takes 200-400ms typical, plus pgvector query overhead, the 500ms budget was being exceeded frequently, silently dropping memory recall. Agent typing delay is already 2-15s humanized, so a 2s recall budget is well within UX tolerance and gives ~4-5x margin over typical embedding latency. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|---|---|---|
| .. | ||
| captain | ||
| cloudflare | ||
| companies | ||
| contacts | ||
| enterprise | ||
| internal | ||
| llm | ||
| messages | ||
| sla | ||
| twilio | ||
| voice | ||
| page_crawler_service.rb | ||