A couple of posts back I walked through the agent and tools we built on top of the Laravel AI SDK for There There, the helpdesk we’re putting together at Spatie. I ended that post on a small cliffhanger. We hand the agent off to a streaming service that turns the SDK’s response stream into events the frontend can consume. That’s this post.
A quick reminder on our angle. After two decades of running our own customer support, we wanted a helpdesk where AI makes support agents faster, not one that tries to replace them. The human reads, thinks, and directs. The model drafts and retrieves in real time while they watch. There There is in private beta right now, and you can apply for early access at there-there.app.
Streaming from Laravel to React
The obvious way to hand an LLM response to the browser is to wait for it to finish, then render the whole thing. That’s easy, but it makes the agent feel sluggish. What we actually want is for the reply to appear word by word, and for the agent’s own reasoning (like which tool it’s calling) to be narrated along the way.
The Laravel AI SDK exposes each step of a conversation as a typed event. We iterate over that stream and emit newline-delimited JSON for the frontend, one line per event.
Here’s the core of our AgentStreamingService.
public function stream(Agent&HasTools $agent, string $message, AgentChat $chat): StreamedResponse { return response()->stream(function () use ($agent, $message, $chat): Generator { yield from $this->streamAgent($agent, $message, $chat); }, 200, [ 'Content-Type' => 'application/x-ndjson', 'Cache-Control' => 'no-cache', ]); } private function streamAgent(Agent&HasTools $agent, string $message, AgentChat $chat): Generator { set_time_limit(120); $fullContent = ''; $toolCalls = []; foreach ($agent->stream($message) as $event) { if ($event instanceof ToolCall) { yield json_encode(['type' => 'tool_call', 'tool_name' => $event->toolCall->name])."n"; } if ($event instanceof TextDelta) { $fullContent .= $event->delta; yield json_encode(['type' => 'delta', 'content' => $event->delta])."n"; } if ($event instanceof ToolResult) { $toolCalls[] = [ 'name' => $event->toolResult->name, 'arguments' => $event->toolResult->arguments, 'result' => $event->toolResult->result, ]; } } $chat->messages()->create([ 'role' => MessageRole::Assistant, 'content' => $fullContent, 'tool_calls' => $toolCalls ?: null, ]); yield json_encode([ 'type' => 'done', 'html' => $this->parser->buildFinalHtml($fullContent), ])."n"; }
A few things are worth calling out. response()->stream() accepts a generator and flushes each yield to the client immediately, which is the whole trick for server-pushed progress. We type the events (delta, tool_call, done) so the frontend can tell a text chunk apart from a tool invocation. And we persist the finished message to the database in the same pass. The stream itself is the source of truth, and the persisted copy is what we rehydrate the next time the user opens the chat.
On the frontend we consume the stream with plain fetch. No library needed, no SSE framing to worry about. Here’s the hook we use, trimmed to the bits that matter.
const response = await fetch(sendUrl, { method: 'POST', headers: { ...csrfHeaders(), 'Content-Type': 'application/json' }, body: JSON.stringify(buildBody(text, html)), signal: controller.signal, }); const reader = response.body.getReader(); const decoder = new TextDecoder(); let buffer = ''; while (true) { const { done, value } = await reader.read(); if (done) break; buffer += decoder.decode(value, { stream: true }); const lines = buffer.split('n'); buffer = lines.pop() ?? ''; for (const line of lines) { processLine(line); } } if (buffer.trim()) { processLine(buffer); }
The buffer pattern is the only thing you have to get right. A single reader.read() call can hand you a partial line, a whole line, or several lines at once. We split on n, keep the last (possibly incomplete) chunk in buffer, and flush it at the end. Every completed line is one of our events.
Dispatching each event to state is a short switch.
if (event.type === 'tool_call') { setMessages((prev) => updateLastStreaming(prev, { toolStatus: formatToolName(event.tool_name), })); } else if (event.type === 'delta') { setMessages((prev) => updateLastStreaming(prev, (last) => ({ content: last.content + event.content, toolStatus: null, }))); } else if (event.type === 'done') { setMessages((prev) => updateLastStreaming(prev, { html: event.html, isStreaming: false, })); }
When a tool_call arrives, we show a small “Looking up tickets” label above the still-streaming reply. When deltas start landing, we clear the tool label and append to the message content. When done arrives, we swap in the server-rendered HTML and flip the streaming flag off. The whole round-trip feels like the agent is typing in front of you.
TODO: short video of the Ask There There agent replying to a question, with a “Looking up tickets” label appearing briefly and then the tokens streaming in.
In closing
There are heavier ways to do this. Server-sent events, websockets, a full streaming library on each side. For a chat interface, none of it is necessary. A generator on the server, NDJSON on the wire, and a small buffer loop on the client is about a hundred lines end to end, and it behaves exactly like a proper streaming LLM chat should.
You can read more about the events the SDK emits in the Laravel AI SDK repo. And if you’d like to try There There yourself, we’re in private beta right now and you can apply for early access at there-there.app.