Skip to content
All work

2026· frontend· ai

Nivram — BrandDrive's AI assistant

An AI assistant that doesn't just answer questions about a business's books — it proposes actions against them, and I built the layer that decides when those are safe to run.

Scope

I built Nivram's client on both platforms. On the web the AI interaction layer is entirely mine. On mobile I built both halves: the streaming shell — socket handler, frame scheduler, chunk merge, message list, chat screens, at 88–100% of surviving lines and 100% of the frame scheduler — and the interaction layer the stream feeds, covering the component registry (84%), the structured parser (96%), the PIN gate (83%), the personal/business resolver (75%), and the hidden message that reports a result back to the model (100%). Of the seventy action cards, sixty-one are majority mine — 78% of the lines across them; colleagues wrote the personal bill-pay and transfer cards. I did not build the apps around all this: BrandDrive mobile alone is twelve people and roughly three thousand commits, of which 109 are mine.

Screens

  • The web client, where the interaction layer is entirely mine. This one frame is the shape the whole study is about: the assistant collects what it needs, renders a summary you can check, then mounts a single button that calls the same production endpoint the app's own customer screen calls.
  • Mobile, at the start. The chips are the assistant's four verbs — create, analyse, compare, track — so a first-time user has somewhere to begin other than a blank prompt, and the remaining quota is shown before it is spent rather than after.
  • Mid-stream. 'Stewing…' is the thinking event of the socket protocol — start, thinking, chunk, end — and the prose beneath it is written out of a partial buffer at one render per frame rather than one per token. Nothing pressable has mounted yet: the text streams, the structure waits until the payload is whole.
  • Finished. Markdown as it arrives — headings, emphasis, nested lists — with the message controls attaching once the answer is complete.
  • Conversation management. A chat begins as a local draft and is promoted to a server thread mid-reply, so the key identifying it changes while a message is still arriving — which is why the socket handler reads the current key instead of closing over it.

Problem

Nivram answers questions about a business's own books — receivables ratios, top suppliers, revenue forecasts, when to restock. Those answers take real time to generate, so a request-response interface leaves an owner staring at a spinner wondering whether it has hung. Streaming fixes that, and React Native does not stream: its `fetch` gives you no readable response body, so the pattern every web tutorial assumes simply does not exist on the platform.

Constraints

  • React Native's `fetch` has no streamable response body — `getReader()` does not exist, so the standard web approach is unavailable.
  • The app already had a shared WebSocket hook, and it stored the last event in `useState`. At Nivram's chunk rate, events overwrite each other and text goes missing.
  • A chat begins as a local draft and is promoted to a server thread mid-conversation — the key identifying the conversation changes while a reply is still arriving.
  • Answers are not prose. The envelope carries a typed payload that drives roughly ninety review cards — create a customer, raise an invoice, pay a bill, move money between wallets — each calling the same production endpoints as the app's own screens, with a transaction PIN in front of anything that spends. The renderer has to hold a half-arrived object whose finished form can move somebody's money.
  • The model ships on its own schedule and the app ships through two app stores. A user on last month's build will be offered actions their copy has never heard of, and there is no way to force them to update.

Decisions

Sockets, because the platform ruled out the alternative

With no streamable response body in React Native, the server emits discrete events over Socket.IO — start, thinking, chunk, end — rather than a chunked HTTP body. Nivram attaches its own listener instead of using the app's shared WebSocket hook, because that hook keeps only the latest event in state and would drop chunks the moment they arrive faster than React re-renders.

One render per frame, not one per token

Tokens arrive far faster than a screen refreshes. Each one is coalesced into a ref and flushed to the store on a `requestAnimationFrame` tick, so the list re-renders about sixty times a second regardless of whether sixty tokens arrived or six hundred. Dispatching per token is the obvious implementation and it makes the UI stutter exactly when the user is watching it most closely.

A listener that survives the conversation being renamed

When a draft is promoted to a real thread the chat key changes underneath an in-flight reply. Re-binding the socket listener on that change would drop the rest of the message, so the handler is held in a ref and reads the current key rather than closing over it. It is the kind of bug that only appears on the first message of a new chat, which is the first thing any user does.

Merge chunks defensively, in both shapes

Chunks arrive sometimes as deltas to append and sometimes as cumulative snapshots of everything so far. The merge checks whether the incoming chunk already starts with what is held and appends or replaces accordingly, so a change of server behaviour degrades into a redundant assignment instead of duplicated text.

Stream the prose, wait for the structure

The envelope carries a message alongside component and data fields. The text is pulled from the partial buffer and rendered as markdown while it arrives, so the answer reads as it is written; anything the user can press mounts only once the object is complete. Those controls are not decorative — they call the same production endpoints as the app's own screens, and some of them move money. A button rendered from a half-parsed payload is a transfer with a missing field behind a confirm the user has already trusted. Arriving a second late is the correct failure.

A registry the model is allowed to outrun

Component types resolve through a central map, and an unrecognised one renders nothing — a warning in development, silence in production. That sounds like giving up and is the opposite: the assistant releases continuously while the app crawls through two review queues, so a user on last month's build will certainly be offered a card their copy has never heard of. The choice is between a missing card and a crash, and only one of those lets the rest of the answer still be useful.

The PIN gate wraps the action instead of living inside it

Money-moving cards are not mounted and then asked for authorisation. A gate component opens the transaction PIN modal first, merges the PIN into the payload, and only then mounts the real action with the authorisation already present — so the component never exists in an unauthorised state. Putting the check inside each card would have meant seventy chances to forget it; this way there is one, and forgetting is not among the available mistakes.

The result is whispered back into the conversation

When a card succeeds it does not just close. A hidden message goes back up the same socket carrying what happened, so the assistant's next turn knows the customer now exists or the transfer went through. Fire-and-forget would have been one line instead of the round trip — and would leave the model confidently offering to create a record it just created, because nothing ever told it otherwise. The conversation is the state; anything that changes the world has to say so.

Only the newest card is live

Every assistant card except the most recent is disabled. Scrolling up and pressing a confirm from three answers ago would fire an action against a payload the conversation has moved past — on a financial app, most likely a second transfer. The stale ones stay visible as a record of what was offered, and refuse to do anything.

Trade-offs

Gave up

Reusing the app's shared WebSocket abstraction

Bought

Chunks that survive arriving faster than React renders. The house abstraction was correct for every other feature and wrong for this one, which is a harder thing to argue for than writing something new.

Gave up

Up to one frame of latency on every token

Bought

A render rate bounded by the display rather than by the model's output rate. Sixteen milliseconds is invisible; a list re-rendering four hundred times a second is not.

Gave up

Interactive components appearing after the text finishes rather than alongside it

Bought

No possibility of rendering a control from a half-parsed payload. For components that move money, arriving late is the correct failure.

Gave up

A hand-written Android keyboard lift instead of the built-in avoidance view

Bought

A composer that stays above an edge-to-edge IME while the answer streams behind it. iOS keeps the standard behaviour; only Android needed the measurement.

Results

Shipped

App Store · Play Store

in production

Store updates cut

OTA

shipped over the air through EAS Update

Render rate

1 / frame

coalesced from an unbounded token rate

Streaming layer

88 – 100%

surviving lines by git blame across the four core files

Interaction system

75 – 100%

registry, parser, PIN gate, resolver, and the hidden result message

Tests in the module

72+

Jest files under the Nivram tree

Ratios are given as raw counts rather than percentages. A percentage implies a precision that twenty attempts do not have.

What I got wrong

  • BACKGROUNDING IS UNHANDLED at the feature level. The socket reconnects with a fresh token when the app returns to the foreground, and an interrupted stream relies on that reconnection rather than any deliberate resume. It works because Socket.IO retries forever, which is not the same as being designed for.

Stack

  • React Native
  • Expo SDK 55
  • TypeScript
  • Expo Router
  • Socket.IO
  • Redux Toolkit
  • TanStack Query
  • MMKV
  • EAS Build & Update
  • Jest
  • Maestro