Skip to the content.

nlsearch is one endpoint on purpose. A bot does not have to know index names or query DSL; it forwards what the person typed and gets back what happened.

POST /_nl
{"prompt": "<what the user typed>", "session": "<one id per conversation>"}

Keeping the conversation

session is any string you like: the chat id from Slack, Teams or WhatsApp, a row id from your own database, a UUID you keep in the browser. Send the same one on every message of a conversation and leave it out for one-off requests.

That one field is the whole of it. There is nothing to create first, nothing to close afterwards, and no state to hold in your bot.

With a session the model sees the earlier turns, so these work:

the user says what happens
“red shoes under 50” a search
“now only the ones in stock” the same search plus a filter
“sort them by price” the same search plus a sort
“just the top three” the same search plus a size
“delete those” a delete_by_query with that same query
“add two more like the first one” a bulk, copying the shape of that document

Without a session each request starts from nothing, so “only the ones in stock” has no idea what “the ones” were.

What is actually stored

One document per session in .nlsearch-history, holding the whole conversation. Each turn is three strings:

{
  "prompt":  "only the ones in stock",
  "answer":  "{\"action\":\"search\",\"index\":\"products\",\"body\":{...}}",
  "outcome": "4 hits: [id 1] Red Running Shoe, [id 4] Green Summer Sandal, ..."
}

On the next request those turns are replayed to the model as an ordinary chat: your prompt, its own previous answer, then a line saying how that answer went.

outcome is the part that makes it more than transcript. After a search it carries the ids and names of the first few hits, which is how a later “make it 22” knows which document you mean. After a failure it carries the error, so the model can correct itself instead of repeating the mistake.

Every turn is kept, so a conversation can be replayed or audited in full. Only the last ten are replayed to the model, because that is all a context window has room for.

.nlsearch-history is a registered system index. Elasticsearch keeps it out of every wildcard and out of _cat/indices, and refuses writes that do not come from the plugin itself. The model is not allowed to name it either; a plan that targets it is refused.

Working with the history

# one conversation
curl -s localhost:9200/.nlsearch-history/_doc/chat-123?pretty

# just the turns
curl -s localhost:9200/.nlsearch-history/_doc/chat-123 | python3 -c '
import json, sys
for i, t in enumerate(json.load(sys.stdin)["_source"]["turns"], 1):
    print(i, t["prompt"])
    print("   ->", t["outcome"])
'

# which conversations exist
curl -s 'localhost:9200/.nlsearch-history/_search?size=50&_source=updated&pretty'

# forget one, or all of them
curl -s -XDELETE localhost:9200/.nlsearch-history/_doc/chat-123
curl -s -XDELETE localhost:9200/.nlsearch-history

Because it is an ordinary index you can treat it like one: snapshot it, add an ILM policy to expire old conversations, or delete by query on updated to drop anything older than a month.

curl -s -XPOST 'localhost:9200/.nlsearch-history/_delete_by_query' \
  -H 'Content-Type: application/json' \
  -d '{"query": {"range": {"updated": {"lt": "now-30d"}}}}'

Nothing expires on its own, so if you run many short-lived conversations, add something like that on a schedule.

Rules of thumb for a bot

Showing the answer

This is the part that makes a bot easy to write. Every answer carries response: a sentence the model wrote from what Elasticsearch returned, ready to send straight to the person.

{
  "action": "search",
  "response": "There are 5 products out of stock.",
  "result": { "hits": { "total": { "value": 5 } } }
}

So the simplest possible bot is: forward the message, print response. No formatting of hits, no counting, no deciding whether a number or a list is the answer. The model already did that.

result is the ordinary Elasticsearch response, there when you want to do more than print a line: the hits are at result.hits.hits, a grouped question at result.aggregations, a write at result.result (created, updated, deleted). Show the sentence, draw a table from the data.

Ask for less when you want less:

response you get when
missing, or both the sentence and the data the usual case for a bot
explain the sentence only the answer goes straight to a person
raw the data only your own code reads it, and you want the speed

Writing the sentence is one more call to the model, so raw is the quickest of the three. The top of the answer is unchanged in every mode: action, index, body and the rest, which is good for a “here is what I did” line or a debug view.

When action is reply, nothing ran and text is a sentence to show the user as is: the request was not about the data, or something was missing.

Errors come back as normal Elasticsearch errors with a reason written for a person, so a bot can show it: 400 when the plan was refused, 404 for an index that is not there, 502 when the model is unreachable or produced nonsense.

Keeping users safe


Back to the start