【问题标题】:Trying to get the frequencies of a .wav file in Python试图在 Python 中获取 .wav 文件的频率
【发布时间】:2019-07-03 20:44:08
【问题描述】:

我知道关于 Python 中 .wav 文件的问题几乎被打死了,但我非常沮丧,因为似乎没有人的答案对我有用。我想做的事情对我来说似乎相对简单:我想确切地知道 .wav 文件中在给定时间的频率。我想知道,比如“从时间n毫秒到n+10毫秒,声音的平均频率是x赫兹”。我看到人们谈论傅立叶变换和 Goertzel 算法,以及各种模块,我似乎无法弄清楚如何去做我所描述的事情。我已经尝试查找诸如“在 python 中查找 wav 文件的频率”之类的内容大约 20 次,但均无济于事。有人可以帮我吗?

我正在寻找的是一个像这个伪代码这样的解决方案,或者至少一个可以做类似伪代码的解决方案:

import some_module_that_can_help_me_do_this as freq

file = 'output.wav'
start_time = 1000  # Start 1000 milliseconds into the file
end_time = 1010  # End 10 milliseconds thereafter

print("Average frequency = " + str(freq.average(start_time, end_time)) + " hz")

请假设(我相信你可以说)我是个数学白痴。这是我的第一个问题,所以要温柔

【问题讨论】:

  • 你的意思是主频吗?请注意,任何真实的音频信号都包含各种频率的混合。您可以简单地计算平均值(例如,来自@Primusa 共享的链接),但我怀疑它是您正在寻找的。​​span>
  • 是的,这听起来像我要找的。对于我的output.wav 文件,我使用的是自己在卡西欧上演奏的录音,所以录音中会有噪音。这就是我想要的主导频率吗?
  • 正确。我的代码对你有用吗?

标签: python python-3.x audio wav


【解决方案1】:

这个答案很晚,但你可以试试这个:

(注意:我不应该为此获得任何荣誉,因为我从其他 SO 帖子和这篇关于使用 Python 进行 FFT 的精彩文章中获得了大部分内容:https://realpython.com/python-scipy-fft/

import numpy as np
from scipy.fft import *
from scipy.io import wavfile


def freq(file, start_time, end_time):

    # Open the file and convert to mono
    sr, data = wavfile.read(file)
    if data.ndim > 1:
        data = data[:, 0]
    else:
        pass

    # Return a slice of the data from start_time to end_time
    dataToRead = data[int(start_time * sr / 1000) : int(end_time * sr / 1000) + 1]

    # Fourier Transform
    N = len(dataToRead)
    yf = rfft(dataToRead)
    xf = rfftfreq(N, 1 / sr)

    # Uncomment these to see the frequency spectrum as a plot
    # plt.plot(xf, np.abs(yf))
    # plt.show()

    # Get the most dominant frequency and return it
    idx = np.argmax(np.abs(yf))
    freq = xf[idx]
    return freq

此代码适用于任何.wav 文件,但它可能略有偏差,因为它只返回最主要的频率,而且它只使用音频的第一个通道(如果不是单声道)。

如果您想了解有关傅里叶变换如何工作的更多信息,请观看 3blue1brown 提供的视频,并提供直观的解释:https://www.youtube.com/watch?v=spUNpyF58BY

【讨论】:

  • 这对我有用,谢谢。
【解决方案2】:

我感到 OP 很沮丧 - 如果有人需要,不应该很难找到如何获取频谱图的值而不是查看频谱图图像:

#!/usr/bin/env python

import librosa
import sys
import numpy as np
import matplotlib.pyplot as plt
import librosa.display

np.set_printoptions(threshold=sys.maxsize)

filename = 'filename.wav'
Fs = 44100
clip, sample_rate = librosa.load(filename, sr=Fs)

n_fft = 1024  # frame length 
start = 0 

hop_length=512

#commented out code to display Spectrogram
X = librosa.stft(clip, n_fft=n_fft, hop_length=hop_length)
#Xdb = librosa.amplitude_to_db(abs(X))
#plt.figure(figsize=(14, 5))
#librosa.display.specshow(Xdb, sr=Fs, x_axis='time', y_axis='hz') 
#If to pring log of frequencies  
#librosa.display.specshow(Xdb, sr=Fs, x_axis='time', y_axis='log')
#plt.colorbar()

#librosa.display.waveplot(clip, sr=Fs)
#plt.show()

#now print all values 

t_samples = np.arange(clip.shape[0]) / Fs
t_frames = np.arange(X.shape[1]) * hop_length / Fs
#f_hertz = np.arange(N / 2 + 1) * Fs / N       # Works only when N is even
f_hertz = np.fft.rfftfreq(n_fft, 1 / Fs)         # Works also when N is odd

#example
print('Time (seconds) of last sample:', t_samples[-1])
print('Time (seconds) of last frame: ', t_frames[-1])
print('Frequency (Hz) of last bin:   ', f_hertz[-1])

print('Time (seconds) :', len(t_samples))

#prints array of time frames 
print('Time of frames (seconds) : ', t_frames)
#prints array of frequency bins
print('Frequency (Hz) : ', f_hertz)

print('Number of frames : ', len(t_frames))
print('Number of bins : ', len(f_hertz))

#This code is working to printout frame by frame intensity of each frequency
#on top line gives freq bins
curLine = 'Bins,'
for b in range(1, len(f_hertz)):
    curLine += str(f_hertz[b]) + ','
print(curLine)

curLine = ''
for f in range(1, len(t_frames)):
    curLine = str(t_frames[f]) + ','
    for b in range(1, len(f_hertz)): #for each frame, we get list of bin values printed
        curLine += str("%.02f" % np.abs(X[b, f])) + ','
        #remove format of the float for full details if needed
        #curLine += str(np.abs(X[b, f])) + ','
        #print other useful info like phase of frequency bin b at frame f.
        #curLine += str("%.02f" % np.angle(X[b, f])) + ',' 
    print(curLine)

【讨论】:

    【解决方案3】:

    如果您想检测pitch 的声音(看起来您确实这样做了),那么就 Python 库而言,您最好的选择是 aubio。请参考此example 进行实施。

    import sys
    from aubio import source, pitch
    
    win_s = 4096
    hop_s = 512 
    
    s = source(your_file, samplerate, hop_s)
    samplerate = s.samplerate
    
    tolerance = 0.8
    
    pitch_o = pitch("yin", win_s, hop_s, samplerate)
    pitch_o.set_unit("midi")
    pitch_o.set_tolerance(tolerance)
    
    pitches = []
    confidences = []
    
    total_frames = 0
    while True:
        samples, read = s()
        pitch = pitch_o(samples)[0]
        pitches += [pitch]
        confidence = pitch_o.get_confidence()
        confidences += [confidence]
        total_frames += read
        if read < hop_s: break
    
    print("Average frequency = " + str(np.array(pitches).mean()) + " hz")
    

    请务必查看docs 了解音高检测方法。

    我还认为您可能对估计平均频率和其他一些音频参数感兴趣,而无需使用任何特殊库。让我们只使用numpy!这应该让您更好地了解如何计算此类音频特征。它基于来自seewave 包的specprop。检查文档以了解计算特征的含义。

    import numpy as np
    
    def spectral_properties(y: np.ndarray, fs: int) -> dict:
        spec = np.abs(np.fft.rfft(y))
        freq = np.fft.rfftfreq(len(y), d=1 / fs)
        spec = np.abs(spec)
        amp = spec / spec.sum()
        mean = (freq * amp).sum()
        sd = np.sqrt(np.sum(amp * ((freq - mean) ** 2)))
        amp_cumsum = np.cumsum(amp)
        median = freq[len(amp_cumsum[amp_cumsum <= 0.5]) + 1]
        mode = freq[amp.argmax()]
        Q25 = freq[len(amp_cumsum[amp_cumsum <= 0.25]) + 1]
        Q75 = freq[len(amp_cumsum[amp_cumsum <= 0.75]) + 1]
        IQR = Q75 - Q25
        z = amp - amp.mean()
        w = amp.std()
        skew = ((z ** 3).sum() / (len(spec) - 1)) / w ** 3
        kurt = ((z ** 4).sum() / (len(spec) - 1)) / w ** 4
    
        result_d = {
            'mean': mean,
            'sd': sd,
            'median': median,
            'mode': mode,
            'Q25': Q25,
            'Q75': Q75,
            'IQR': IQR,
            'skew': skew,
            'kurt': kurt
        }
    
        return result_d
    

    【讨论】:

    • 工作得很好,但我不得不调整一些东西。它首先给了我一个错误,说 confidence 没有定义,我不知道如何计算它,所以我只是将它硬编码为 0.8。然后它给了我一个连接类型错误,因为np.array(pitches).mean() 没有返回一个字符串,所以我str()ed 它现在它工作得很好。谢谢
    • 感谢@kitty4537 指出问题的努力。固定的。作为一个额外的东西,我在代码中添加了带你计算各种光谱特征的代码。享受吧!
    【解决方案4】:

    尝试以下方法,它对我有用,我生成了一个频率为 1234 的正弦波文件 from this page.

    from scipy.io import wavfile
    
    def freq(file, start_time, end_time):
        sample_rate, data = wavfile.read(file)
        start_point = int(sample_rate * start_time / 1000)
        end_point = int(sample_rate * end_time / 1000)
        length = (end_time - start_time) / 1000
        counter = 0
        for i in range(start_point, end_point):
            if data[i] < 0 and data[i+1] > 0:
                counter += 1
        return counter/length    
    
    freq("sin.wav", 1000 ,2100)
    1231.8181818181818
    

    已编辑:稍微清理一下 for 循环

    【讨论】:

    • 谢谢,但似乎有问题。当我在这里运行代码时(使用freq 调用的适当打印语句),它给了我错误ValueError: The truth value of an array with more than one element is ambiguous. Use a.any() or a.all()。所以我尝试将第 10 行更改为 if data[i].all() &lt; 0 and data[i+1].all() &gt; 0: 然后 if data[i].any() &lt; 0 and data[i+1].any() &gt; 0: ,但无论我为开始/结束指定什么时间,这两次我都得到了 0.0
    • 很奇怪,你能分享一下你传递给这个函数的文件和开始/结束时间吗?
    • 你运行的是哪个版本的 Python?
    • 运行 Python 3.7。这是引发ValueError 的版本:https://pastebin.com/aT4F4bg4 这是一个无论我使用什么开始/结束时间都返回0.0 的版本:https://pastebin.com/CB9TwbLT 这是一个也只返回0.0https://pastebin.com/rCSw5Rc4跨度>
    • 这种算法只能在信号中存在单一频率的情况下工作。本质上,这是一种计算过零率的简化方法。
    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 2021-03-18
    • 1970-01-01
    • 2022-01-03
    • 2016-01-20
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多