Technology

72697 readers

2495 users here now

This is a most excellent place for technology news and articles.

Our Rules

Follow the lemmy.world rules.
Only tech related news or articles.
Be excellent to each other!
Mod approved content bots can post up to 10 articles per day.
Threads asking for personal tech support may be deleted.
Politics threads may be removed.
No memes allowed as posts, OK to post as comments.
Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
Check for duplicates before posting, duplicates may be removed
Accounts 7 days and younger will have their posts automatically removed.

Approved Bots

founded 2 years ago

MODERATORS

L3s@lemmy.world

enu@lemmy.world

technopagan@lemmy.world

L4s@lemmy.world

L3s@hackingne.ws

L4s@hackingne.ws

295

Microsoft’s VASA-1 can deepfake a person with one photo and one audio track (arstechnica.com)

submitted 1 year ago by return2ozma@lemmy.world to c/technology@lemmy.world

70 comments fedilink hide all child comments

(page 2) 22 comments

sorted by: hot top controversial new old

[–] Dasus@lemmy.world 2 points 1 year ago* (last edited 1 year ago) (3 children)

One use of this I'm in favour of is recreating Majel Barret's voice as an AI for computer systems.

load more comments (3 replies)

[–] autotldr@lemmings.world 2 points 1 year ago

This is the best summary I could come up with:

On Tuesday, Microsoft Research Asia unveiled VASA-1, an AI model that can create a synchronized animated video of a person talking or singing from a single photo and an existing audio track.

In the future, it could power virtual avatars that render locally and don't require video feeds—or allow anyone with similar tools to take a photo of a person found online and make them appear to say whatever they want.

To show off the model, Microsoft created a VASA-1 research page featuring many sample videos of the tool in action, including people singing and speaking in sync with pre-recorded audio tracks.

The examples also include some more fanciful generations, such as Mona Lisa rapping to an audio track of Anne Hathaway performing a "Paparazzi" song on Conan O'Brien.

While the Microsoft researchers tout potential positive applications like enhancing educational equity, improving accessibility, and providing therapeutic companionship, the technology could also easily be misused.

"We are opposed to any behavior to create misleading or harmful contents of real persons, and are interested in applying our technique for advancing forgery detection," write the researchers.

The original article contains 797 words, the summary contains 183 words. Saved 77%. I'm a bot and I'm open source!

load more comments