
Rob Strechay
2
APPEARANCES
2
PODCASTS
012
DEC 30
JAN 6
JAN 13
JAN 20
JAN 27
FEB 3
FEB 10
FEB 17
FEB 24
MAR 3
MAR 10
MAR 17
MAR 24
MAR 31
APR 7
APR 14
APR 21
APR 28
MAY 5
MAY 12
MAY 19
MAY 26
JUN 2
JUN 9
JUN 16
JUN 23
JUN 30
JUL 7
JUL 14
JUL 21
JUL 28
AUG 4
AUG 11
AUG 18
AUG 25
SEP 1
SEP 8
SEP 15
SEP 22
SEP 29
OCT 6
OCT 13
OCT 20
OCT 27
NOV 3
NOV 10
NOV 17
NOV 24
DEC 1
DEC 8
DEC 15
DEC 22
DEC 29
JAN 5
JAN 12
JAN 19
JAN 26
FEB 2
FEB 9
FEB 16
FEB 23
MAR 2
MAR 9
MAR 16
MAR 23
MAR 30
APR 6
APR 13
APR 20
APR 27
MAY 4
MAY 11
MAY 18
MAY 25
JUN 1
JUN 8
JUN 15
JUN 22
JUN 29
JUL 6
JUL 13
JUL 20
JUL 27
AUG 3
AUG 10
AUG 17
AUG 24
AUG 31
SEP 7
SEP 14
Sep 11, 2026
The State of Enterprise AI (Rob Strechay—VentureBeat)
9:34
10:10
10:27
11:01
11:39

Rob StrechayGUEST
very expensive very inefficient very context uh rich but context broad and i think when you want it to give you specific answers without hallucinations shrinking that that whole context window down and i don't mean the real context window i mean its frame of reference for that model will help it really focus in on certain things and kind of keep it corralled a little bit so i agree with you you're going to have i think the orchestration layer will be more broad layer of models.

Rob StrechayGUEST
And I think that when you start to get down into the task oriented agents and agentic processes and tools, those will be SLMs that are tuned or shrunk or distilled or what have you.

Rob StrechayGUEST
Like if I look at it and go, if I'm building out a new company or if I'm, retrofitting a retail company and i need quick answers when i'm sitting in a retail location i'm in a shopping store i'm you know supermarket or something like that i may want to distill a model down and have very specific to all the skews in here what aisles they're in and have all that tracked pretty locally to that retail location And have those answers over the Wi-Fi and tell people if you connect to our Wi-Fi and use the app here, it'll give you all this.

Rob StrechayGUEST
And by the way, then I can also use tools like location, you know, centric guiding to get people to their exact location of where are the gherkin pickles in this particular store.

Matt HousleyHOST
I think that's very connected to what's happening at enterprises as well that have to spend a lot of money on tokens with OpenAI or Google or other companies.
GPU Hoarding is Over. The $401B Reality Check
3:36
3:46
4:31
4:49
10:55
11:19

Rob StrechayGUEST
Yeah, I mean, I think, again, when they looked at it, the cost per inference, and like I talked about before, had jumped from 34% to 41%.

Rob StrechayGUEST
And when you started to look at how organizations were saying that they really needed more control and security over this ai 72 of them admitted that they had a lack of sufficient control over it and that they were really moving towards inference and trying to help understand what was going on and they're looking at different ways of building it as well with about 11 percent moving up to 17 percent where they're actually controlling their full stack from an infrastructure on up perspective because that is the best way that they see fit of being able to focus on that TCO and ROAI of the inference.

Rob StrechayGUEST
And inference really, again, in talking to others, you know, the large neoclouds and hyperscalers, they're already seeing, you know, 60, 80% of the workloads happening on their infrastructure have really shifted to inference from training.

Rob StrechayGUEST
And I think, again, most people have figured out that they're not in the training game and they're trying to weigh that battle between being a token consumer and a token producer and how they solve for that to actually meet their numbers.
6 MINS LATER

Erran BergerGUEST
You need to know how the business metrics that product, AI product is driving is going to move you know revenue and then and then you need to look at the opportunity cost of investing that compute on something else um as well and so so all that requires you know instrumentation it requires rigor engineering rigor uh and discipline

Rob StrechayGUEST
yeah i just to add to that i mean i think from the data the data is backing that up right i mean people are getting a little bit gun shy of just going and buying and hoarding gpus in fact you know folks were looking at if it fell from the 20s down into the teens from a perspective of that from 20.8 to 15.4 and it's kind of a retreat from that what we had been seeing and i think also part of that and I'm interested in this too, is just the amount of memory and flash and the cost that's risen.