SlideShare ist ein Scribd-Unternehmen logo
1 von 19
Nikolay V. Karpov (nkarpov(а)hse.ru)

Duration
 1 module, 10 weeks, 40 academic hours


Requirements
 3 practical works at home using Java, Matlab
  or others (lms.hse.ru)
 Final assessment
   2 Modelling Speech Production Acoustics
   3 Time/Frequency Representation.
    Properties of Digital Filters
   4 Linear Predictive Modelling
   5-6 Speech Coding
   7 Phonetics
   8 Speech Synthesis
   9-10 Speech Recognition
   Lingvocourse.ru
    http://lingvocourse.ru/wiki/index.php/Speech_recog
    nition
   Digital speech processing, synthesis, and recognition
    / Sadaoki Furui.- 2nd ed.,
   Speech Analysis Synthesis and Perception
    http://hear.ai.uiuc.edu/ECE537/PDF/main-all.pdf
   FUNDAMETALS OF SPEECH RECOGNITION: A SHORT
    COURSE
    http://speech.tifr.res.in/tutorials/fundamentalOfASR_
    picone96.pdf
   Speech Processing. 20 lectures in the Spring Term.
    Mike Brookes
    http://www.ee.ic.ac.uk/hp/staff/dmb/courses/speec
    h/speech.htm
   Coding
   Synthesis
   Recognition
   Identity Verification
   Enhancement
What: To transmit/store a speech waveform using as few bits as
possible while retaining high quality
Why: To save bandwidth in telecoms applications and to reduce
memory storage requirements.
How:
 Correlation ⇒Predictability ⇒Redundancy
    ◦ Predict waveform samples from previous samples and transmit only the
      prediction error
    ◦ Autocorrelation is Fourier transform of power spectrum: a peaky spectrum
      ⇒strong short-term correlations (~ 0.5 ms)
    ◦ Voiced speech is almost periodic ⇒strong long-term correlations (~ 10
      ms)
   Devote few bits to the aspects of speech where errors are least
    noticeable
    ◦ High amplitude speech will mask noise at the same frequency
   Ignore aspects of the speech that are inaudible
    ◦ Power spectrum is much more important than precise waveform
    ◦ For aperiodic sounds, the fine detail of the spectrum does not matter
What: To convert a text string into a speech waveform
Why: For technology to communicate when a display would be
inconvenient because:
 (a) Too big, (b) Eyes busy, (c) Via phone, (d) In the dark, (e)
  Moving around
Problems:
 The spelling of words doesn‟t match their sound
    ◦ Pronunciation rules + an exceptions dictionary
   Some words have multiple meanings + sounds
    ◦ Must guess which is the correct sound
   Simplistic speech models sound mechanical
    ◦ Can use extracts from real speech
   Speech sounds are influenced by adjacent phonemes
    ◦ Use phoneme pairs from real speech
   Important words must be slightly louder
    ◦ Must try to understand the text unit
   Voice pitch and talking speed must vary smoothly throughout a
    sentence
    ◦ Must be able to change pitch and speed without affecting formant
      frequencies
What: To convert a speech waveform into text
Why: To communicate and control technology when a keyboard would be
inconvenient because:
 (a) Too big, (b) Hands busy, (c) Via phone, (d) In the dark, (e) Moving
   around
Problems:
 The spelling of words doesn‟t match their sound
    ◦ Have a big phonetic dictionary
   The waveform of a word varies a lot between different speakers (or even
    the same speaker)
    ◦ Extract features from the speech waveform that are more consistent than the
      waveform
   The extracted features won‟t be exactly repeatable
    ◦ Characterize them with a probability distribution
   Speech sounds are influenced by adjacent phonemes
    ◦ Use context-dependent probability distributions
   Speaking speed varies enormously
    ◦ Try all possible speaking speeds
   No clear boundary between words or phonemes
    ◦ Try all possible boundaries
Speech waves conveys:
  Speaker meaning
  Individual information
  Emotion of speaker


Phrase(sentence) -> word units -> word ->
syllables -> phonemes
  Russian
а э и о у ы п п' б б' м м' ф ф' в в' т т' д д' н н' с
с' з з' р р' л л' ш ж щ җ ц ч й к к' г г' х х„

 English
http://en.wikipedia.org/wiki/English_phonology
   Speakers and listeners divide words into component
    sounds called phonemes.
    ◦ Native speakers agree on the phonemes that make up a
      particular word
    ◦ There are about 42 phonemes in English
 The phonemes in a particular word may vary with
dialect
    ◦ High amplitude speech will mask noise at the same
      frequency
   The actual sound that corresponds to a particular
    phoneme depends on:
    ◦   the adjacent phonemes in the word or sentence
    ◦   the accent of the speaker
    ◦   the talking speed
    ◦   whether it is a formal or informal occasion
   Turbulence: air moving quickly through a
    small hole (e.g./s/ in “size”)
   Explosion: pressure built up behind a
    blockage is suddenly released (e.g. /p/ in
    “pop”)
   Vocal Cords(Fold) Vibration
• airflow through vocal folds (vocal cords) reduces the
pressure and they snap shut (Bernoulli effect)
• muscle tension and air pressure buildup force the folds
open again and the process repeats
• frequency of vibration (fx) determined by tension in
vocal folds and pressure from lungs
• for normal breathing and voiceless sounds (e.g. /s/) the
vocal folds are held wide open and don‟t vibrate
   Vowel /а/, /о/, /у/
   Consonant
    ◦ Unvoiced
      Fricative /ш/, /щ/, /ф/, /х/
      Plosive /п/, /к/, /т/
      Affricate /ч/, /ц/
    ◦ Voiced
        Fricative /ж/, /җ/, /в/, /р/
        Plosive /б/, /г/, /д/
        Diphthongs /oj/
        Nasal /н/, /м/
        Semivowel /r/, /j/, /w/
   The sound spectrum is modified by the shape
    of the vocal tract. This is determined by
    movements of the jaw, tongue and lips.
   The resonant frequencies of the vocal tract
    cause peaks in the spectrum called formants.
   The first two formant frequencies are roughly
    determined by the distances from the tongue
    hump to the larynx and to the lips
    respectively.
Principal characteristics of speech
Principal characteristics of speech
Principal characteristics of speech
Principal characteristics of speech
Principal characteristics of speech

Weitere ähnliche Inhalte

Was ist angesagt?

Phonetics presentation part II
Phonetics presentation   part IIPhonetics presentation   part II
Phonetics presentation part IIShermila Azariah
 
Phonetics presentation at ccnust by Monir Hossen
Phonetics presentation at ccnust by Monir Hossen Phonetics presentation at ccnust by Monir Hossen
Phonetics presentation at ccnust by Monir Hossen Monir Hossen
 
Speech Mechanism
Speech MechanismSpeech Mechanism
Speech Mechanismflattsph
 
Speech considerations for cd/ dentistry dental implants
Speech considerations for cd/ dentistry dental implantsSpeech considerations for cd/ dentistry dental implants
Speech considerations for cd/ dentistry dental implantsIndian dental academy
 
Speech Processes (Phonation and Articulation)
Speech Processes (Phonation and Articulation)Speech Processes (Phonation and Articulation)
Speech Processes (Phonation and Articulation)Christian Sebastian
 
Introduction phonetics
Introduction   phoneticsIntroduction   phonetics
Introduction phoneticsVivine McLeary
 
Csd 210 introduction to phonetics i and ii
Csd 210 introduction to phonetics i and iiCsd 210 introduction to phonetics i and ii
Csd 210 introduction to phonetics i and iiJake Probst
 
Speech organ and manner of articulation
Speech organ and manner of articulationSpeech organ and manner of articulation
Speech organ and manner of articulationYanti95
 
Speech organ uzma
Speech organ uzmaSpeech organ uzma
Speech organ uzmauzma bashir
 
Speech consideration in complete denture
Speech consideration in complete dentureSpeech consideration in complete denture
Speech consideration in complete dentureethan1hunt
 

Was ist angesagt? (20)

Phonetics presentation part II
Phonetics presentation   part IIPhonetics presentation   part II
Phonetics presentation part II
 
Phonetics presentation at ccnust by Monir Hossen
Phonetics presentation at ccnust by Monir Hossen Phonetics presentation at ccnust by Monir Hossen
Phonetics presentation at ccnust by Monir Hossen
 
Speech Mechanism
Speech MechanismSpeech Mechanism
Speech Mechanism
 
Speech considerations for cd/ dentistry dental implants
Speech considerations for cd/ dentistry dental implantsSpeech considerations for cd/ dentistry dental implants
Speech considerations for cd/ dentistry dental implants
 
Speech Processes (Phonation and Articulation)
Speech Processes (Phonation and Articulation)Speech Processes (Phonation and Articulation)
Speech Processes (Phonation and Articulation)
 
Introduction phonetics
Introduction   phoneticsIntroduction   phonetics
Introduction phonetics
 
Csd 210 introduction to phonetics i and ii
Csd 210 introduction to phonetics i and iiCsd 210 introduction to phonetics i and ii
Csd 210 introduction to phonetics i and ii
 
Speech mechanism
Speech mechanismSpeech mechanism
Speech mechanism
 
Phonetics
PhoneticsPhonetics
Phonetics
 
Speech organ and manner of articulation
Speech organ and manner of articulationSpeech organ and manner of articulation
Speech organ and manner of articulation
 
The resonating-parts (1)
The resonating-parts (1)The resonating-parts (1)
The resonating-parts (1)
 
Resonators
ResonatorsResonators
Resonators
 
Consonants
ConsonantsConsonants
Consonants
 
Phonetics report
Phonetics reportPhonetics report
Phonetics report
 
Lecture phonetics
Lecture phoneticsLecture phonetics
Lecture phonetics
 
Speech organ uzma
Speech organ uzmaSpeech organ uzma
Speech organ uzma
 
Affricate sounds 2010
Affricate sounds 2010Affricate sounds 2010
Affricate sounds 2010
 
Speech consideration in complete denture
Speech consideration in complete dentureSpeech consideration in complete denture
Speech consideration in complete denture
 
Consonant
ConsonantConsonant
Consonant
 
Consonant
ConsonantConsonant
Consonant
 

Andere mochten auch

"Automatic speech recognition for mobile applications in Yandex" — Fran Campi...
"Automatic speech recognition for mobile applications in Yandex" — Fran Campi..."Automatic speech recognition for mobile applications in Yandex" — Fran Campi...
"Automatic speech recognition for mobile applications in Yandex" — Fran Campi...Yandex
 
The features of the connected speech final
The features of the connected speech finalThe features of the connected speech final
The features of the connected speech finalHina Honey
 
Blending, phrasing and intonation
Blending, phrasing and intonationBlending, phrasing and intonation
Blending, phrasing and intonationRyan Lualhati
 
Measuring the Effectiveness of the Promotional Program
Measuring the Effectiveness of the Promotional ProgramMeasuring the Effectiveness of the Promotional Program
Measuring the Effectiveness of the Promotional ProgramIndrajit Bage
 
Establishing Objectives and Budgeting for the Promotional Program
Establishing Objectives and Budgeting for the Promotional ProgramEstablishing Objectives and Budgeting for the Promotional Program
Establishing Objectives and Budgeting for the Promotional ProgramIndrajit Bage
 
speech production in psycholinguistics
speech production in psycholinguistics speech production in psycholinguistics
speech production in psycholinguistics Aseel K. Mahmood
 
Chap20 International Advertising And Promotion
Chap20 International Advertising And PromotionChap20 International Advertising And Promotion
Chap20 International Advertising And PromotionPhoenix media & event
 
Comparing Broadsheet and Tabloid newspapers
Comparing Broadsheet and Tabloid newspapersComparing Broadsheet and Tabloid newspapers
Comparing Broadsheet and Tabloid newspapersjodieholmes
 

Andere mochten auch (13)

Chapter21
Chapter21Chapter21
Chapter21
 
"Automatic speech recognition for mobile applications in Yandex" — Fran Campi...
"Automatic speech recognition for mobile applications in Yandex" — Fran Campi..."Automatic speech recognition for mobile applications in Yandex" — Fran Campi...
"Automatic speech recognition for mobile applications in Yandex" — Fran Campi...
 
The features of the connected speech final
The features of the connected speech finalThe features of the connected speech final
The features of the connected speech final
 
Blending, phrasing and intonation
Blending, phrasing and intonationBlending, phrasing and intonation
Blending, phrasing and intonation
 
Measuring the Effectiveness of the Promotional Program
Measuring the Effectiveness of the Promotional ProgramMeasuring the Effectiveness of the Promotional Program
Measuring the Effectiveness of the Promotional Program
 
Connected speech features
Connected speech featuresConnected speech features
Connected speech features
 
Stages of speaking
Stages of speakingStages of speaking
Stages of speaking
 
Establishing Objectives and Budgeting for the Promotional Program
Establishing Objectives and Budgeting for the Promotional ProgramEstablishing Objectives and Budgeting for the Promotional Program
Establishing Objectives and Budgeting for the Promotional Program
 
speech production in psycholinguistics
speech production in psycholinguistics speech production in psycholinguistics
speech production in psycholinguistics
 
Chap20 International Advertising And Promotion
Chap20 International Advertising And PromotionChap20 International Advertising And Promotion
Chap20 International Advertising And Promotion
 
The organs of speech and their function
The organs of speech and their functionThe organs of speech and their function
The organs of speech and their function
 
Comparing Broadsheet and Tabloid newspapers
Comparing Broadsheet and Tabloid newspapersComparing Broadsheet and Tabloid newspapers
Comparing Broadsheet and Tabloid newspapers
 
Properties of Sound
Properties of SoundProperties of Sound
Properties of Sound
 

Ähnlich wie Principal characteristics of speech

Phonetics & phonology, INTRODUCTION, Dr, Salama Embarak
Phonetics & phonology, INTRODUCTION, Dr, Salama EmbarakPhonetics & phonology, INTRODUCTION, Dr, Salama Embarak
Phonetics & phonology, INTRODUCTION, Dr, Salama EmbarakAbdulsalam Mohammed
 
Class 09 emerson_phonetics_fall2014_phonemes_allophones_vot_epg
Class 09 emerson_phonetics_fall2014_phonemes_allophones_vot_epgClass 09 emerson_phonetics_fall2014_phonemes_allophones_vot_epg
Class 09 emerson_phonetics_fall2014_phonemes_allophones_vot_epgLisa Lavoie
 
speech processing basics
speech processing basicsspeech processing basics
speech processing basicssivakumar m
 
speech recognition and removal of disfluencies
speech recognition and removal of disfluenciesspeech recognition and removal of disfluencies
speech recognition and removal of disfluenciesAnkit Sharma
 
Speech signal processing lizy
Speech signal processing lizySpeech signal processing lizy
Speech signal processing lizyLizy Abraham
 
Phonetics full
Phonetics fullPhonetics full
Phonetics fullHina Honey
 
(Emerson) Phonetics & Phonology.pptx
(Emerson) Phonetics & Phonology.pptx(Emerson) Phonetics & Phonology.pptx
(Emerson) Phonetics & Phonology.pptxShamsUlFatah
 
Presentation for China Forum (1).ppt
Presentation for China Forum (1).pptPresentation for China Forum (1).ppt
Presentation for China Forum (1).pptRAJALAKSHMIJ10
 
Ch 9 Language and Speech Processing.pptx
Ch 9 Language and Speech Processing.pptxCh 9 Language and Speech Processing.pptx
Ch 9 Language and Speech Processing.pptxLarry195181
 
Teaching alphabetics and fluency in reading
Teaching alphabetics and fluency in readingTeaching alphabetics and fluency in reading
Teaching alphabetics and fluency in readingMarcia Luptak
 
Speech and Language Processing
Speech and Language ProcessingSpeech and Language Processing
Speech and Language ProcessingVikalp Mahendra
 
Introduction to audiovidual translation by adriana serban
Introduction to audiovidual translation by adriana serbanIntroduction to audiovidual translation by adriana serban
Introduction to audiovidual translation by adriana serbanyiling666
 

Ähnlich wie Principal characteristics of speech (20)

Part1 speech basics
Part1 speech basicsPart1 speech basics
Part1 speech basics
 
Phonetics & phonology, INTRODUCTION, Dr, Salama Embarak
Phonetics & phonology, INTRODUCTION, Dr, Salama EmbarakPhonetics & phonology, INTRODUCTION, Dr, Salama Embarak
Phonetics & phonology, INTRODUCTION, Dr, Salama Embarak
 
Class 09 emerson_phonetics_fall2014_phonemes_allophones_vot_epg
Class 09 emerson_phonetics_fall2014_phonemes_allophones_vot_epgClass 09 emerson_phonetics_fall2014_phonemes_allophones_vot_epg
Class 09 emerson_phonetics_fall2014_phonemes_allophones_vot_epg
 
Isolated English Word Recognition System: Appropriate for Bengali-accented En...
Isolated English Word Recognition System: Appropriate for Bengali-accented En...Isolated English Word Recognition System: Appropriate for Bengali-accented En...
Isolated English Word Recognition System: Appropriate for Bengali-accented En...
 
B110512
B110512B110512
B110512
 
Language
LanguageLanguage
Language
 
speech processing basics
speech processing basicsspeech processing basics
speech processing basics
 
speech recognition and removal of disfluencies
speech recognition and removal of disfluenciesspeech recognition and removal of disfluencies
speech recognition and removal of disfluencies
 
Say That Again? Enhancing Your Accent Acumen
Say That Again? Enhancing Your Accent AcumenSay That Again? Enhancing Your Accent Acumen
Say That Again? Enhancing Your Accent Acumen
 
Speech signal processing lizy
Speech signal processing lizySpeech signal processing lizy
Speech signal processing lizy
 
Week 3 phonology copy
Week 3  phonology   copyWeek 3  phonology   copy
Week 3 phonology copy
 
Phonetics full
Phonetics fullPhonetics full
Phonetics full
 
(Emerson) Phonetics & Phonology.pptx
(Emerson) Phonetics & Phonology.pptx(Emerson) Phonetics & Phonology.pptx
(Emerson) Phonetics & Phonology.pptx
 
Presentation for China Forum (1).ppt
Presentation for China Forum (1).pptPresentation for China Forum (1).ppt
Presentation for China Forum (1).ppt
 
Ch 9 Language and Speech Processing.pptx
Ch 9 Language and Speech Processing.pptxCh 9 Language and Speech Processing.pptx
Ch 9 Language and Speech Processing.pptx
 
Teaching alphabetics and fluency in reading
Teaching alphabetics and fluency in readingTeaching alphabetics and fluency in reading
Teaching alphabetics and fluency in reading
 
Phonology
PhonologyPhonology
Phonology
 
Speech and Language Processing
Speech and Language ProcessingSpeech and Language Processing
Speech and Language Processing
 
Phonetics
PhoneticsPhonetics
Phonetics
 
Introduction to audiovidual translation by adriana serban
Introduction to audiovidual translation by adriana serbanIntroduction to audiovidual translation by adriana serban
Introduction to audiovidual translation by adriana serban
 

Mehr von Nikolay Karpov

Идентификация уровня сложности текста и его адаптация
Идентификация уровня сложности текста и его адаптацияИдентификация уровня сложности текста и его адаптация
Идентификация уровня сложности текста и его адаптацияNikolay Karpov
 
Идентификация уровня ложности текста и его адаптация
Идентификация уровня ложности текста и его адаптацияИдентификация уровня ложности текста и его адаптация
Идентификация уровня ложности текста и его адаптацияNikolay Karpov
 
Теория и практика обработки естественного языка
Теория и практика обработки естественного языкаТеория и практика обработки естественного языка
Теория и практика обработки естественного языкаNikolay Karpov
 
Speech waves in tube and filters
Speech waves in tube and filtersSpeech waves in tube and filters
Speech waves in tube and filtersNikolay Karpov
 
Speech signal time frequency representation
Speech signal time frequency representationSpeech signal time frequency representation
Speech signal time frequency representationNikolay Karpov
 

Mehr von Nikolay Karpov (8)

Идентификация уровня сложности текста и его адаптация
Идентификация уровня сложности текста и его адаптацияИдентификация уровня сложности текста и его адаптация
Идентификация уровня сложности текста и его адаптация
 
Идентификация уровня ложности текста и его адаптация
Идентификация уровня ложности текста и его адаптацияИдентификация уровня ложности текста и его адаптация
Идентификация уровня ложности текста и его адаптация
 
Cepstral coefficients
Cepstral coefficientsCepstral coefficients
Cepstral coefficients
 
Теория и практика обработки естественного языка
Теория и практика обработки естественного языкаТеория и практика обработки естественного языка
Теория и практика обработки естественного языка
 
Linear prediction
Linear predictionLinear prediction
Linear prediction
 
Speech waves in tube and filters
Speech waves in tube and filtersSpeech waves in tube and filters
Speech waves in tube and filters
 
Speech signal time frequency representation
Speech signal time frequency representationSpeech signal time frequency representation
Speech signal time frequency representation
 
Tagger numbers
Tagger numbersTagger numbers
Tagger numbers
 

Kürzlich hochgeladen

How to do quick user assign in kanban in Odoo 17 ERP
How to do quick user assign in kanban in Odoo 17 ERPHow to do quick user assign in kanban in Odoo 17 ERP
How to do quick user assign in kanban in Odoo 17 ERPCeline George
 
4.16.24 Poverty and Precarity--Desmond.pptx
4.16.24 Poverty and Precarity--Desmond.pptx4.16.24 Poverty and Precarity--Desmond.pptx
4.16.24 Poverty and Precarity--Desmond.pptxmary850239
 
ROLES IN A STAGE PRODUCTION in arts.pptx
ROLES IN A STAGE PRODUCTION in arts.pptxROLES IN A STAGE PRODUCTION in arts.pptx
ROLES IN A STAGE PRODUCTION in arts.pptxVanesaIglesias10
 
Keynote by Prof. Wurzer at Nordex about IP-design
Keynote by Prof. Wurzer at Nordex about IP-designKeynote by Prof. Wurzer at Nordex about IP-design
Keynote by Prof. Wurzer at Nordex about IP-designMIPLM
 
ANG SEKTOR NG agrikultura.pptx QUARTER 4
ANG SEKTOR NG agrikultura.pptx QUARTER 4ANG SEKTOR NG agrikultura.pptx QUARTER 4
ANG SEKTOR NG agrikultura.pptx QUARTER 4MiaBumagat1
 
Virtual-Orientation-on-the-Administration-of-NATG12-NATG6-and-ELLNA.pdf
Virtual-Orientation-on-the-Administration-of-NATG12-NATG6-and-ELLNA.pdfVirtual-Orientation-on-the-Administration-of-NATG12-NATG6-and-ELLNA.pdf
Virtual-Orientation-on-the-Administration-of-NATG12-NATG6-and-ELLNA.pdfErwinPantujan2
 
ICS2208 Lecture6 Notes for SL spaces.pdf
ICS2208 Lecture6 Notes for SL spaces.pdfICS2208 Lecture6 Notes for SL spaces.pdf
ICS2208 Lecture6 Notes for SL spaces.pdfVanessa Camilleri
 
Measures of Position DECILES for ungrouped data
Measures of Position DECILES for ungrouped dataMeasures of Position DECILES for ungrouped data
Measures of Position DECILES for ungrouped dataBabyAnnMotar
 
4.18.24 Movement Legacies, Reflection, and Review.pptx
4.18.24 Movement Legacies, Reflection, and Review.pptx4.18.24 Movement Legacies, Reflection, and Review.pptx
4.18.24 Movement Legacies, Reflection, and Review.pptxmary850239
 
Millenials and Fillennials (Ethical Challenge and Responses).pptx
Millenials and Fillennials (Ethical Challenge and Responses).pptxMillenials and Fillennials (Ethical Challenge and Responses).pptx
Millenials and Fillennials (Ethical Challenge and Responses).pptxJanEmmanBrigoli
 
Activity 2-unit 2-update 2024. English translation
Activity 2-unit 2-update 2024. English translationActivity 2-unit 2-update 2024. English translation
Activity 2-unit 2-update 2024. English translationRosabel UA
 
How to Add Barcode on PDF Report in Odoo 17
How to Add Barcode on PDF Report in Odoo 17How to Add Barcode on PDF Report in Odoo 17
How to Add Barcode on PDF Report in Odoo 17Celine George
 
Dust Of Snow By Robert Frost Class-X English CBSE
Dust Of Snow By Robert Frost Class-X English CBSEDust Of Snow By Robert Frost Class-X English CBSE
Dust Of Snow By Robert Frost Class-X English CBSEaurabinda banchhor
 
HỌC TỐT TIẾNG ANH 11 THEO CHƯƠNG TRÌNH GLOBAL SUCCESS ĐÁP ÁN CHI TIẾT - CẢ NĂ...
HỌC TỐT TIẾNG ANH 11 THEO CHƯƠNG TRÌNH GLOBAL SUCCESS ĐÁP ÁN CHI TIẾT - CẢ NĂ...HỌC TỐT TIẾNG ANH 11 THEO CHƯƠNG TRÌNH GLOBAL SUCCESS ĐÁP ÁN CHI TIẾT - CẢ NĂ...
HỌC TỐT TIẾNG ANH 11 THEO CHƯƠNG TRÌNH GLOBAL SUCCESS ĐÁP ÁN CHI TIẾT - CẢ NĂ...Nguyen Thanh Tu Collection
 
Expanded definition: technical and operational
Expanded definition: technical and operationalExpanded definition: technical and operational
Expanded definition: technical and operationalssuser3e220a
 
Active Learning Strategies (in short ALS).pdf
Active Learning Strategies (in short ALS).pdfActive Learning Strategies (in short ALS).pdf
Active Learning Strategies (in short ALS).pdfPatidar M
 
The Contemporary World: The Globalization of World Politics
The Contemporary World: The Globalization of World PoliticsThe Contemporary World: The Globalization of World Politics
The Contemporary World: The Globalization of World PoliticsRommel Regala
 
TEACHER REFLECTION FORM (NEW SET........).docx
TEACHER REFLECTION FORM (NEW SET........).docxTEACHER REFLECTION FORM (NEW SET........).docx
TEACHER REFLECTION FORM (NEW SET........).docxruthvilladarez
 

Kürzlich hochgeladen (20)

How to do quick user assign in kanban in Odoo 17 ERP
How to do quick user assign in kanban in Odoo 17 ERPHow to do quick user assign in kanban in Odoo 17 ERP
How to do quick user assign in kanban in Odoo 17 ERP
 
YOUVE GOT EMAIL_FINALS_EL_DORADO_2024.pptx
YOUVE GOT EMAIL_FINALS_EL_DORADO_2024.pptxYOUVE GOT EMAIL_FINALS_EL_DORADO_2024.pptx
YOUVE GOT EMAIL_FINALS_EL_DORADO_2024.pptx
 
INCLUSIVE EDUCATION PRACTICES FOR TEACHERS AND TRAINERS.pptx
INCLUSIVE EDUCATION PRACTICES FOR TEACHERS AND TRAINERS.pptxINCLUSIVE EDUCATION PRACTICES FOR TEACHERS AND TRAINERS.pptx
INCLUSIVE EDUCATION PRACTICES FOR TEACHERS AND TRAINERS.pptx
 
4.16.24 Poverty and Precarity--Desmond.pptx
4.16.24 Poverty and Precarity--Desmond.pptx4.16.24 Poverty and Precarity--Desmond.pptx
4.16.24 Poverty and Precarity--Desmond.pptx
 
ROLES IN A STAGE PRODUCTION in arts.pptx
ROLES IN A STAGE PRODUCTION in arts.pptxROLES IN A STAGE PRODUCTION in arts.pptx
ROLES IN A STAGE PRODUCTION in arts.pptx
 
Keynote by Prof. Wurzer at Nordex about IP-design
Keynote by Prof. Wurzer at Nordex about IP-designKeynote by Prof. Wurzer at Nordex about IP-design
Keynote by Prof. Wurzer at Nordex about IP-design
 
ANG SEKTOR NG agrikultura.pptx QUARTER 4
ANG SEKTOR NG agrikultura.pptx QUARTER 4ANG SEKTOR NG agrikultura.pptx QUARTER 4
ANG SEKTOR NG agrikultura.pptx QUARTER 4
 
Virtual-Orientation-on-the-Administration-of-NATG12-NATG6-and-ELLNA.pdf
Virtual-Orientation-on-the-Administration-of-NATG12-NATG6-and-ELLNA.pdfVirtual-Orientation-on-the-Administration-of-NATG12-NATG6-and-ELLNA.pdf
Virtual-Orientation-on-the-Administration-of-NATG12-NATG6-and-ELLNA.pdf
 
ICS2208 Lecture6 Notes for SL spaces.pdf
ICS2208 Lecture6 Notes for SL spaces.pdfICS2208 Lecture6 Notes for SL spaces.pdf
ICS2208 Lecture6 Notes for SL spaces.pdf
 
Measures of Position DECILES for ungrouped data
Measures of Position DECILES for ungrouped dataMeasures of Position DECILES for ungrouped data
Measures of Position DECILES for ungrouped data
 
4.18.24 Movement Legacies, Reflection, and Review.pptx
4.18.24 Movement Legacies, Reflection, and Review.pptx4.18.24 Movement Legacies, Reflection, and Review.pptx
4.18.24 Movement Legacies, Reflection, and Review.pptx
 
Millenials and Fillennials (Ethical Challenge and Responses).pptx
Millenials and Fillennials (Ethical Challenge and Responses).pptxMillenials and Fillennials (Ethical Challenge and Responses).pptx
Millenials and Fillennials (Ethical Challenge and Responses).pptx
 
Activity 2-unit 2-update 2024. English translation
Activity 2-unit 2-update 2024. English translationActivity 2-unit 2-update 2024. English translation
Activity 2-unit 2-update 2024. English translation
 
How to Add Barcode on PDF Report in Odoo 17
How to Add Barcode on PDF Report in Odoo 17How to Add Barcode on PDF Report in Odoo 17
How to Add Barcode on PDF Report in Odoo 17
 
Dust Of Snow By Robert Frost Class-X English CBSE
Dust Of Snow By Robert Frost Class-X English CBSEDust Of Snow By Robert Frost Class-X English CBSE
Dust Of Snow By Robert Frost Class-X English CBSE
 
HỌC TỐT TIẾNG ANH 11 THEO CHƯƠNG TRÌNH GLOBAL SUCCESS ĐÁP ÁN CHI TIẾT - CẢ NĂ...
HỌC TỐT TIẾNG ANH 11 THEO CHƯƠNG TRÌNH GLOBAL SUCCESS ĐÁP ÁN CHI TIẾT - CẢ NĂ...HỌC TỐT TIẾNG ANH 11 THEO CHƯƠNG TRÌNH GLOBAL SUCCESS ĐÁP ÁN CHI TIẾT - CẢ NĂ...
HỌC TỐT TIẾNG ANH 11 THEO CHƯƠNG TRÌNH GLOBAL SUCCESS ĐÁP ÁN CHI TIẾT - CẢ NĂ...
 
Expanded definition: technical and operational
Expanded definition: technical and operationalExpanded definition: technical and operational
Expanded definition: technical and operational
 
Active Learning Strategies (in short ALS).pdf
Active Learning Strategies (in short ALS).pdfActive Learning Strategies (in short ALS).pdf
Active Learning Strategies (in short ALS).pdf
 
The Contemporary World: The Globalization of World Politics
The Contemporary World: The Globalization of World PoliticsThe Contemporary World: The Globalization of World Politics
The Contemporary World: The Globalization of World Politics
 
TEACHER REFLECTION FORM (NEW SET........).docx
TEACHER REFLECTION FORM (NEW SET........).docxTEACHER REFLECTION FORM (NEW SET........).docx
TEACHER REFLECTION FORM (NEW SET........).docx
 

Principal characteristics of speech

  • 1. Nikolay V. Karpov (nkarpov(а)hse.ru) Duration  1 module, 10 weeks, 40 academic hours Requirements  3 practical works at home using Java, Matlab or others (lms.hse.ru)  Final assessment
  • 2. 2 Modelling Speech Production Acoustics  3 Time/Frequency Representation. Properties of Digital Filters  4 Linear Predictive Modelling  5-6 Speech Coding  7 Phonetics  8 Speech Synthesis  9-10 Speech Recognition
  • 3. Lingvocourse.ru http://lingvocourse.ru/wiki/index.php/Speech_recog nition  Digital speech processing, synthesis, and recognition / Sadaoki Furui.- 2nd ed.,  Speech Analysis Synthesis and Perception http://hear.ai.uiuc.edu/ECE537/PDF/main-all.pdf  FUNDAMETALS OF SPEECH RECOGNITION: A SHORT COURSE http://speech.tifr.res.in/tutorials/fundamentalOfASR_ picone96.pdf  Speech Processing. 20 lectures in the Spring Term. Mike Brookes http://www.ee.ic.ac.uk/hp/staff/dmb/courses/speec h/speech.htm
  • 4. Coding  Synthesis  Recognition  Identity Verification  Enhancement
  • 5. What: To transmit/store a speech waveform using as few bits as possible while retaining high quality Why: To save bandwidth in telecoms applications and to reduce memory storage requirements. How:  Correlation ⇒Predictability ⇒Redundancy ◦ Predict waveform samples from previous samples and transmit only the prediction error ◦ Autocorrelation is Fourier transform of power spectrum: a peaky spectrum ⇒strong short-term correlations (~ 0.5 ms) ◦ Voiced speech is almost periodic ⇒strong long-term correlations (~ 10 ms)  Devote few bits to the aspects of speech where errors are least noticeable ◦ High amplitude speech will mask noise at the same frequency  Ignore aspects of the speech that are inaudible ◦ Power spectrum is much more important than precise waveform ◦ For aperiodic sounds, the fine detail of the spectrum does not matter
  • 6. What: To convert a text string into a speech waveform Why: For technology to communicate when a display would be inconvenient because:  (a) Too big, (b) Eyes busy, (c) Via phone, (d) In the dark, (e) Moving around Problems:  The spelling of words doesn‟t match their sound ◦ Pronunciation rules + an exceptions dictionary  Some words have multiple meanings + sounds ◦ Must guess which is the correct sound  Simplistic speech models sound mechanical ◦ Can use extracts from real speech  Speech sounds are influenced by adjacent phonemes ◦ Use phoneme pairs from real speech  Important words must be slightly louder ◦ Must try to understand the text unit  Voice pitch and talking speed must vary smoothly throughout a sentence ◦ Must be able to change pitch and speed without affecting formant frequencies
  • 7. What: To convert a speech waveform into text Why: To communicate and control technology when a keyboard would be inconvenient because:  (a) Too big, (b) Hands busy, (c) Via phone, (d) In the dark, (e) Moving around Problems:  The spelling of words doesn‟t match their sound ◦ Have a big phonetic dictionary  The waveform of a word varies a lot between different speakers (or even the same speaker) ◦ Extract features from the speech waveform that are more consistent than the waveform  The extracted features won‟t be exactly repeatable ◦ Characterize them with a probability distribution  Speech sounds are influenced by adjacent phonemes ◦ Use context-dependent probability distributions  Speaking speed varies enormously ◦ Try all possible speaking speeds  No clear boundary between words or phonemes ◦ Try all possible boundaries
  • 8. Speech waves conveys:  Speaker meaning  Individual information  Emotion of speaker Phrase(sentence) -> word units -> word -> syllables -> phonemes
  • 9.  Russian а э и о у ы п п' б б' м м' ф ф' в в' т т' д д' н н' с с' з з' р р' л л' ш ж щ җ ц ч й к к' г г' х х„  English http://en.wikipedia.org/wiki/English_phonology
  • 10. Speakers and listeners divide words into component sounds called phonemes. ◦ Native speakers agree on the phonemes that make up a particular word ◦ There are about 42 phonemes in English  The phonemes in a particular word may vary with dialect ◦ High amplitude speech will mask noise at the same frequency  The actual sound that corresponds to a particular phoneme depends on: ◦ the adjacent phonemes in the word or sentence ◦ the accent of the speaker ◦ the talking speed ◦ whether it is a formal or informal occasion
  • 11.
  • 12. Turbulence: air moving quickly through a small hole (e.g./s/ in “size”)  Explosion: pressure built up behind a blockage is suddenly released (e.g. /p/ in “pop”)  Vocal Cords(Fold) Vibration • airflow through vocal folds (vocal cords) reduces the pressure and they snap shut (Bernoulli effect) • muscle tension and air pressure buildup force the folds open again and the process repeats • frequency of vibration (fx) determined by tension in vocal folds and pressure from lungs • for normal breathing and voiceless sounds (e.g. /s/) the vocal folds are held wide open and don‟t vibrate
  • 13. Vowel /а/, /о/, /у/  Consonant ◦ Unvoiced  Fricative /ш/, /щ/, /ф/, /х/  Plosive /п/, /к/, /т/  Affricate /ч/, /ц/ ◦ Voiced  Fricative /ж/, /җ/, /в/, /р/  Plosive /б/, /г/, /д/  Diphthongs /oj/  Nasal /н/, /м/  Semivowel /r/, /j/, /w/
  • 14. The sound spectrum is modified by the shape of the vocal tract. This is determined by movements of the jaw, tongue and lips.  The resonant frequencies of the vocal tract cause peaks in the spectrum called formants.  The first two formant frequencies are roughly determined by the distances from the tongue hump to the larynx and to the lips respectively.