Skip to main content
Rob Strechay

Rob Strechay

Sep 11, 2026

9:34
expensive
9:34
very expensive very inefficient very context uh rich but context broad and i think when you want it to give you specific answers without hallucinations shrinking that that whole context window down and i don't mean the real context window i mean its frame of reference for that model will help it really focus in on certain things and kind of keep it corralled a little bit so i agree with you you're going to have i think the orchestration layer will be more broad layer of models.
10:10
And I think that when you start to get down into the task oriented agents and agentic processes and tools, those will be SLMs that are tuned or shrunk or distilled or what have you.
10:26
So they can fit someplace.
10:27
Like if I look at it and go, if I'm building out a new company or if I'm, retrofitting a retail company and i need quick answers when i'm sitting in a retail location i'm in a shopping store i'm you know supermarket or something like that i may want to distill a model down and have very specific to all the skews in here what aisles they're in and have all that tracked pretty locally to that retail location And have those answers over the Wi-Fi and tell people if you connect to our Wi-Fi and use the app here, it'll give you all this.
11:01
And by the way, then I can also use tools like location, you know, centric guiding to get people to their exact location of where are the gherkin pickles in this particular store.
11:14
And I can't find them.
11:39
I think that's very connected to what's happening at enterprises as well that have to spend a lot of money on tokens with OpenAI or Google or other companies.
3:34
Can you give us a breakdown of what the data is saying?
3:36
Yeah, I mean, I think, again, when they looked at it, the cost per inference, and like I talked about before, had jumped from 34% to 41%.
3:46
And when you started to look at how organizations were saying that they really needed more control and security over this ai 72 of them admitted that they had a lack of sufficient control over it and that they were really moving towards inference and trying to help understand what was going on and they're looking at different ways of building it as well with about 11 percent moving up to 17 percent where they're actually controlling their full stack from an infrastructure on up perspective because that is the best way that they see fit of being able to focus on that TCO and ROAI of the inference.
4:31
And inference really, again, in talking to others, you know, the large neoclouds and hyperscalers, they're already seeing, you know, 60, 80% of the workloads happening on their infrastructure have really shifted to inference from training.
4:49
And I think, again, most people have figured out that they're not in the training game and they're trying to weigh that battle between being a token consumer and a token producer and how they solve for that to actually meet their numbers.

6 MINS LATER

10:55
You need to know how the business metrics that product, AI product is driving is going to move you know revenue and then and then you need to look at the opportunity cost of investing that compute on something else um as well and so so all that requires you know instrumentation it requires rigor engineering rigor uh and discipline
11:19
yeah i just to add to that i mean i think from the data the data is backing that up right i mean people are getting a little bit gun shy of just going and buying and hoarding gpus in fact you know folks were looking at if it fell from the 20s down into the teens from a perspective of that from 20.8 to 15.4 and it's kind of a retreat from that what we had been seeing and i think also part of that and I'm interested in this too, is just the amount of memory and flash and the cost that's risen.
11:55
Some people were smart enough to hoard that early on.

We value your privacy

We use cookies to understand how you use our platform and to improve your experience. Click “Accept All” to consent, or “Decline non-essential” to opt out of non-essential cookies. Read our Privacy Policy.