And, like, it was back in 2013 when you started, and since 2016 we saw a surge in deep learning that evolution and, and so in speech technology also has evolved.
So now in the current era, if I ask you, like, given how deep fake models, uh, keep advancing, like the...
As attackers are building better models that are be- that are getting better and better at generating deep fakes.
Also at the same time, there's other side where the detectors are building better detection systems and racing to catch up.
So I believe that's kind of like the, uh, like, like you mentioned, uh, eternal arms race in any kind of like security related topic.
Um, yeah, so, so definitely the speech synthesis technologies are quite different from, uh, to, today compared to 15, 20 years ago.
And, well, when it comes to, let's say, developing something like, uh, countermeasures or different solutions, uh, so I, I think, I think this development is something that we have to kind of like keep in mind all the time.
And, well, in, in the context of this ASVS spoof challenges in particular, like we have, uh, kind of like designed, uh, these challenges actually to, in a way that, well, you, you are supposed to, to build this kind of like, uh, real versus fake detector.
Given a few seconds of speech, give a confidence score, how, how likely is it to be, to, uh, be real versus, uh, something like, uh, deep, deep fake.
And, uh, I, I think in the design of these, uh, challenges right from the beginning, I, I think the design philosophy was that, okay, we, uh, need to keep the training material something like, uh, to consist of, let's say, particular speech synthesis techniques and particular set of speakers and material.
But at the evolution set, uh, we, we have done something purposefully different one, so because okay, that's, it's just to mimic the, the kind of like the, this, this kind of like adversarial situation that if you are, if you are a hacker, uh, likely you wouldn't be telling in advance, uh, the technique that you are gonna use for, let's say, uh, fooling a biometric system.
So it's, in this sense, I think the, this task is quite challenging actually to address because we, unlike in many other classification tasks where we could safely say that, "Okay, here is the training distribution, uh, here is the test distribution," and they are quite similar and, you know, identically distributed speech materials at the training and testing.
But in this kind of like adversarial set ups, I believe it's kind of like built up in the, in the task almost that, okay, it's something, uh, unexpected, uh, will happen at, in the real world and, okay, we just try to build this kind of like proxy when developing this kind of like test bench to develop solution.
To the arms race, well, I, I per- personally, uh, believe because, okay, the, the both technologies, the both, uh, let's say synthesis technologies and the defense technologies are moving forward and I, I believe there will be an eventual limit beyond which we, you know, uh, this kind of like problem set up that we have, like the real or fake detection.
There is gonna be limit to it, and I, I think also therefore it's also important to consider other solutions like watermarking and, um, well, maybe even in the legal/ethical guidelines and all of that kind of like really complex stuff.
And, like, it was back in 2013 when you started, and since 2016 we saw a surge in deep learning that evolution and, and so in speech technology also has evolved.
So now in the current era, if I ask you, like, given how deep fake models, uh, keep advancing, like the...
As attackers are building better models that are be- that are getting better and better at generating deep fakes.
Also at the same time, there's other side where the detectors are building better detection systems and racing to catch up.
So I believe that's kind of like the, uh, like, like you mentioned, uh, eternal arms race in any kind of like security related topic.
Um, yeah, so, so definitely the speech synthesis technologies are quite different from, uh, to, today compared to 15, 20 years ago.
And, well, when it comes to, let's say, developing something like, uh, countermeasures or different solutions, uh, so I, I think, I think this development is something that we have to kind of like keep in mind all the time.
And, well, in, in the context of this ASVS spoof challenges in particular, like we have, uh, kind of like designed, uh, these challenges actually to, in a way that, well, you, you are supposed to, to build this kind of like, uh, real versus fake detector.
Given a few seconds of speech, give a confidence score, how, how likely is it to be, to, uh, be real versus, uh, something like, uh, deep, deep fake.
And, uh, I, I think in the design of these, uh, challenges right from the beginning, I, I think the design philosophy was that, okay, we, uh, need to keep the training material something like, uh, to consist of, let's say, particular speech synthesis techniques and particular set of speakers and material.
But at the evolution set, uh, we, we have done something purposefully different one, so because okay, that's, it's just to mimic the, the kind of like the, this, this kind of like adversarial situation that if you are, if you are a hacker, uh, likely you wouldn't be telling in advance, uh, the technique that you are gonna use for, let's say, uh, fooling a biometric system.
So it's, in this sense, I think the, this task is quite challenging actually to address because we, unlike in many other classification tasks where we could safely say that, "Okay, here is the training distribution, uh, here is the test distribution," and they are quite similar and, you know, identically distributed speech materials at the training and testing.
But in this kind of like adversarial set ups, I believe it's kind of like built up in the, in the task almost that, okay, it's something, uh, unexpected, uh, will happen at, in the real world and, okay, we just try to build this kind of like proxy when developing this kind of like test bench to develop solution.
To the arms race, well, I, I per- personally, uh, believe because, okay, the, the both technologies, the both, uh, let's say synthesis technologies and the defense technologies are moving forward and I, I believe there will be an eventual limit beyond which we, you know, uh, this kind of like problem set up that we have, like the real or fake detection.
There is gonna be limit to it, and I, I think also therefore it's also important to consider other solutions like watermarking and, um, well, maybe even in the legal/ethical guidelines and all of that kind of like really complex stuff.
The rest of this transcript — segmented and speaker-labeled, so you land on the exact moment something was said
Search every transcript — by keyword, by phrase, or by meaning, across every show Radar indexes
Trends — what is surging across podcasts, measured against its own baseline
Alerts — when a name you follow appears in a newly indexed episode
No account is needed to search Radar.