US Patent 11749285 Speech transcription using multiple data sources

Patent 11749285 was granted and assigned to Meta on September, 2023 by the United States Patent and Trademark Office.

Overview Structured Data Issues Contributors Activity

All edits

Edits on 7 Sep, 2023

"Created via: Entity Importer"

Golden AI

created this topic on 7 Sep, 2023

Edits made to:

Infobox (+29 properties)

Article (+937 characters)

‌

US Patent 11749285 Speech transcription using multiple data sources

Article

This disclosure describes transcribing speech using audio, image, and other data. A system is described that includes an audio capture system configured to capture audio data associated with a plurality of speakers, an image capture system configured to capture images of one or more of the plurality of speakers, and a speech processing engine. The speech processing engine may be configured to recognize a plurality of speech segments in the audio data, identify, for each speech segment of the plurality of speech segments and based on the images, a speaker associated with the speech segment, transcribe each of the plurality of speech segments to produce a transcription of the plurality of speech segments including, for each speech segment in the plurality of speech segments, an indication of the speaker associated with the speech segment, and analyze the transcription to produce additional data derived from the transcription.

Infobox

Is a