| | |
| | | ```sh |
| | | cd egs/aishell/paraformer |
| | | ``` |
| | | |
| | | Then you can directly start the recipe as follows: |
| | | ```sh |
| | | conda activate funasr |
| | | . ./run.sh |
| | | ``` |
| | | The training log files are saved in `exp/*_train_*/log/train.log.*`, which can be viewed using the following command: |
| | | |
| | | The training log files are saved in `${exp_dir}/exp/${model_dir}/log/train.log.*`, which can be viewed using the following command: |
| | | ```sh |
| | | vim exp/*_train_*/log/train.log.0 |
| | | ``` |
| | | |
| | | Users can observe the training loss, prediction accuracy and other training information, like follows: |
| | | ```text |
| | | ... 1epoch:train:751-800batch:800num_updates: ... loss_ctc=106.703, loss_att=86.877, acc=0.029, loss_pre=1.552 ... |
| | | ... 1epoch:train:801-850batch:850num_updates: ... loss_ctc=107.890, loss_att=87.832, acc=0.029, loss_pre=1.702 ... |
| | | ``` |
| | | |
| | | Also, users can use tensorboard to observe these training information by the following command: |
| | | ```sh |
| | | tensorboard --logdir ${exp_dir}/exp/${model_dir}/tensorboard/train |
| | | ``` |
| | | |
| | | At the end of each epoch, the evaluation metrics are calculated on the validation set, like follows: |
| | | ```text |
| | | ... [valid] loss_ctc=99.914, cer_ctc=1.000, loss_att=80.512, acc=0.029, cer=0.971, wer=1.000, loss_pre=1.952, loss=88.285 ... |
| | | ``` |
| | | |
| | | The inference results are saved in `exp/*_train_*/decode_asr_*/$dset`. The main two files are `text.cer` and `text.cer.txt`. `text.cer` saves the comparison between the recognized text and the reference text, like follows: |
| | | The inference results are saved in `${exp_dir}/exp/${model_dir}/decode_asr_*/$dset`. The main two files are `text.cer` and `text.cer.txt`. `text.cer` saves the comparison between the recognized text and the reference text, like follows: |
| | | ```text |
| | | ... |
| | | BAC009S0764W0213(nwords=11,cor=11,ins=0,del=0,sub=0) corr=100.00%,cer=0.00% |
| | |
| | | - `CUDA_VISIBLE_DEVICES`: `0,1` (Default), visible gpu list |
| | | - `gpu_num`: `2` (Default), the number of GPUs used for training |
| | | - `gpu_inference`: `true` (Default), whether to use GPUs for decoding |
| | | - `njob`: `1` (Default), for CPU decoding, indicating the total number of CPU jobs; for GPU decoding, indicating the number of jobs on each GPU |
| | | - `njob`: `1` (Default),for CPU decoding, indicating the total number of CPU jobs; for GPU decoding, indicating the number of jobs on each GPU |
| | | - `raw_data`: the raw path of AISHELL-1 dataset |
| | | - `feats_dir`: the path for saving processed data |
| | | - `token_type`: `char` (Default), indicate how to process text |
| | | - `type`: `sound` (Default), set the input type |
| | | - `scp`: `wav.scp` (Default), set the input file |
| | | - `nj`: `64` (Default), the number of jobs for data preparation |
| | | - `speed_perturb`: `"0.9, 1.0 ,1.1"` (Default), the range of speech perturbed |
| | | - `exp_dir`: the path for saving experimental results |
| | |
| | | |
| | | ### Stage 2: Dictionary Preparation |
| | | This stage processes the dictionary, which is used as a mapping between label characters and integer indices during ASR training. The processed dictionary file is saved as `$feats_dir/data/$lang_toekn_list/$token_type/tokens.txt`. An example of `tokens.txt` is as follows: |
| | | * `tokens.txt` |
| | | ``` |
| | | <blank> |
| | | <s> |
| | |
| | | 龟 |
| | | <unk> |
| | | ``` |
| | | * `<blank>`: indicates the blank token for CTC |
| | | * `<s>`: indicates the start-of-sentence token |
| | | * `</s>`: indicates the end-of-sentence token |
| | | * `<unk>`: indicates the out-of-vocabulary token |
| | | * `<blank>`: indicates the blank token for CTC, must be in the first line |
| | | * `<s>`: indicates the start-of-sentence token, must be in the second line |
| | | * `</s>`: indicates the end-of-sentence token, must be in the third line |
| | | * `<unk>`: indicates the out-of-vocabulary token, must be in the last line |
| | | |
| | | ### Stage 3: LM Training |
| | | |
| | |
| | | |
| | | We support two parameters to specify the training steps, namely `max_epoch` and `max_update`. `max_epoch` indicates the total training epochs while `max_update` indicates the total training steps. If these two parameters are specified at the same time, once the training reaches any one of these two parameters, the training will be stopped. |
| | | |
| | | * Tensorboard |
| | | |
| | | Users can use tensorboard to observe the loss, learning rate, etc. Please run the following command: |
| | | ``` |
| | | tensorboard --logdir ${exp_dir}/exp/${model_dir}/tensorboard/train |
| | | ``` |
| | | |
| | | ### Stage 5: Decoding |
| | | This stage generates the recognition results and calculates the `CER` to verify the performance of the trained model. |
| | | |
| | |
| | | * Performance |
| | | |
| | | We adopt `CER` to verify the performance. The results are in `$exp_dir/exp/$model_dir/$decoding_yaml_name/$average_model_name/$dset`, namely `text.cer` and `text.cer.txt`. `text.cer` saves the comparison between the recognized text and the reference text while `text.cer.txt` saves the final `CER` results. The following is an example of `text.cer`: |
| | | * `text.cer` |
| | | ``` |
| | | ... |
| | | BAC009S0764W0213(nwords=11,cor=11,ins=0,del=0,sub=0) corr=100.00%,cer=0.00% |