Skip to main content
Whisper

Whisper

SoftwareWikipedia

Search complete. 129 mentions across 24 episodes found for "Whisper".

Sep 15, 2026

Adam SchwabGUEST
39:04
But the manual stuff actually, and getting the data right in the right format and fixing the mistakes takes longer than actual prompting, bizarrely.
Lisa TehHOST
39:15
I need to get you onto Whisper, remember? I'm always like, Adam, use Whisper.
Lisa TehHOST
39:19
Especially because, yeah, listen to how fast you talk.
Lisa TehHOST
39:21
I mean, you talk, you'll be able to like, it might be like, cannot compute talking way too fast.
Pavankumar Reddy MuddireddyGUEST
14:36
So we have the trunk, which is a three B, uh, text model that we train, uh, that we call Mistral series of models, and audio input is provided to the model through a audio encoder.
Pavankumar Reddy MuddireddyGUEST
14:52
But unlike, say, Whisper or models like that, where the audio input goes through an encoder, which is then fed into the decoder through a cross attention, here audio, uh, encoder produces, uh, s- tokens, um, in this case, uh, embed continuous representations through embeddings.
Pavankumar Reddy MuddireddyGUEST
15:13
And then they're fed into the main decoder model, uh, just as a direct token input, similar to how you would feed text input.
Pavankumar Reddy MuddireddyGUEST
15:23
Uh, in, in the text case, it's a rather simple encoding scheme.
Pavankumar Reddy MuddireddyGUEST
15:26
You, uh, s- uh, you send it through a tokenizer and you get token IDs, and then you just have a embedding table.
Pavankumar Reddy MuddireddyGUEST
15:34
Uh, in this case, the encoder is a little bit more sophisticated.
Pavankumar Reddy MuddireddyGUEST
15:38
Uh, at least in the VoxelChat, the encoder is very close to Whisper encoder, although for the l-later models, we optimized it and adapted it.
Pavankumar Reddy MuddireddyGUEST
15:47
Uh, we just tried to reduce the number of layers to the, uh, minimum number required for getting the performance.
Thomas SteinerGUEST
59:59
One of the reasons we have the imperative API as well is that you can do negotiating.
Thomas SteinerGUEST
60:04
So let's say you built a transcription app and you can work well with Whisper Tiny, like the smallest model in the Whisper family.
Thomas SteinerGUEST
60:12
But then of course, if the user already locally has Whisper massive, whatever, giant in their cost cache, you wouldn't say no, right? You wouldn't download a tiny model if they already have the big one.
Thomas SteinerGUEST
60:23
So that's also some legit cases where you would just want to allow probing and just say like, yeah, this is a pretty good use case where we say, yeah, I know that I work with the smallest model, but if I have a bigger model already available, I will just take that.
Pete WardenGUEST
33:45
they're actually often pretty good.
Pete WardenGUEST
33:50
Like Quen, 1.7 billion, I think it is, actually does a better job than Whisper, significantly better job on some languages like Mandarin.
Pete WardenGUEST
34:06
And that is a, and, you know, the same goes for Gemma.
Pete WardenGUEST
34:13
The trouble is that as general purpose models, like one of the failure cases with Gemma is that you upload audio of some speech.
JamieHOST
1:49
In local model developments, Desert Ant Labs debuted 18 on-device models covering transcription, audio cleanup, personal data redaction, language identification and video clipping, built to run entirely on-device with no cloud roundtrip and no token costs, free up to 100,000 monthly active devices.
JamieHOST
2:08
Its transcription model turns 10 minutes of audio into text in two seconds on an iPhone, nearly five times faster than Whisper, and its redaction model catches most sensitive strings in a 12-megabyte footprint.
JamieHOST
2:20
These are exactly the utility features consumer apps default to a cloud model call for today, and a free, faster, offline alternative quietly erodes a slice of per-token API revenue.
JamieHOST
2:31
DeepSeek released v4, OneFlash, a 552 billion parameter multimodal model under an MIT license, claiming near parity with its own larger v4 Pro base while activating only a quarter of the parameters and cutting per-token cache memory to under a kilobyte.
Leo LaporteHOST
4:12
The first thing I noticed immediately is the dictation, thankfully, is decent.
Leo LaporteHOST
4:16
It's still not as good as the state-of-the-art, the Whisper dictation, for instance, that, uh, Paris uses.
Leo LaporteHOST
4:22
But it's, but it's a lot better.
Paris MartineauHOST
4:25
Whisper dictation is so good.
Matthew CassinelliGUEST
4:26
Oh, yeah.
Leo LaporteHOST
4:27
It's amazing.
Matthew CassinelliGUEST
4:38
I use stuff like Monologue on desktop.
Leo LaporteHOST
4:39
And because Apple's such a walled garden, uh, Apple's such a walled garden, you can't...
Leo LaporteHOST
4:12
The first thing I noticed immediately is the dictation thankfully is decent.
Leo LaporteHOST
4:16
It's still not as good as the state-of-the-art, the Whisper dictation, for instance, that, uh, Paris uses, but it's, but it's a lot better.
Paris MartineauHOST
4:25
Whisper dictation is so good.
Matthew CassinelliGUEST
4:26
Oh, yeah.
Leo LaporteHOST
4:27
It's amazing.
Matthew CassinelliGUEST
4:39
... stuff like Monologue on the desktop
Leo LaporteHOST
4:39
... such a walled garden, uh, Apple's such a walled garden, you can't...
Leo LaporteHOST
4:43
I mean, on all, on my computers I can use Whisper AI, but I, I, I can't really use it on the iPhone.
Edo SegalHOST
1:01
En el frente del software empresarial, Daniel tiene dos lanzamientos que llegaron de madrugada.
DanielCORRESPONDENT
1:06
Microsoft lanzó en silencio MI Transcribe 2 a diez centavos por hora, un movimiento de presión directa sobre los precios de Whisper y Assembly AI.
DanielCORRESPONDENT
1:16
Y xAI abrió Grok Bot a equipos empresariales, añadiendo controles de política y barreras de seguridad para despliegue autónomo.
DanielCORRESPONDENT
1:24
El primer nivel formal de producto empresarial de xAI.
DanielCORRESPONDENT
3:46
Ahora cubre sesenta idiomas, frente a cuarenta y tres en junio, y añade diarización de hablantes, estilos configurables y marcas de tiempo a nivel de palabra.
DanielCORRESPONDENT
3:56
Microsoft afirma que ocupa el primer lugar en el benchmark Fluers, ejerciendo presión directa sobre OpenAI Whisper, Google y Eleven Labs en el mercado de transcripción empresarial.
Edo SegalHOST
4:07
Cinco historias, un hilo conductor.
Edo SegalHOST
4:10
La infraestructura de la IA diplomática, de hardware, de modelos y de precios está siendo renegociada en cada capa de manera simultánea.
Laura OsborneGUEST
13:46
No, I think it's been a lot of conversations at the moment have been really purpose driven.
Laura OsborneGUEST
13:52
So it's things like Whisper comes up a lot because people are specifically looking for transcription and translation and those kinds of storylines, which is, you know, if you... search for what do I use with transcription? What's a good model? Wisp is obviously going to be the first one that comes up.
Laura OsborneGUEST
14:09
So I do think there's still a bit more of an education piece to be done around this.
Laura OsborneGUEST
14:13
And I think it's evolved a lot.

14 more episodes mention Whisper.

Create an account to see the whole feed, search across every transcript, and follow the entities you care about.

We value your privacy

We use cookies to understand how you use our platform and to improve your experience. Click “Accept All” to consent, or “Decline non-essential” to opt out of non-essential cookies. Read our Privacy Policy.