AI in the stackLesson 1 of 47 min

A model is just another API

Strip the mystique and it is a slow, expensive, unreliable third-party call.

Architecturally, calling a language model is calling someone else’s HTTP endpoint. It has an address, it takes a body, it returns a body, it can fail, it can rate limit you. Everything you learned in the APIs course applies unchanged.

What differs is the shape of the numbers. A normal API call takes tens of milliseconds; a model call takes seconds. A normal API call costs effectively nothing; a model call costs real money per request. And a normal API returns the same answer for the same input, while a model may not.

The consequences of "slow and expensive"
  1. Seconds, not milliseconds
    Stream the response or the page appears frozen.
  2. Costs per call
    A retry loop is now a bill. Cap and monitor it.
  3. Can be unavailable
    It is a dependency. Decide what your app does when it is down.
  4. Non-deterministic
    The same input can produce different output. Tests need to account for that.

What to remember

  • A model call is an ordinary third-party API call with unusual cost and latency.
  • The key belongs on the server, always.
  • Streaming improves perceived speed, not real speed.

Terms in this lesson

Show field notes toggles a search param the loader reads. With it off, the slow promise is never created, so nothing streams.