You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
About Dataset: The Flickr 8k Audio Caption Corpus contains 40,000 spoken captions of 8,000 natural images. It was collected in 2015 to investigate multimodal learning schemes for unsupervised speech pattern discovery. This corpus only includes audio recordings, and not the original text captions or associated images.
The text was updated successfully, but these errors were encountered:
Claim Dataset: Flickr Audio Caption
About Dataset: The Flickr 8k Audio Caption Corpus contains 40,000 spoken captions of 8,000 natural images. It was collected in 2015 to investigate multimodal learning schemes for unsupervised speech pattern discovery. This corpus only includes audio recordings, and not the original text captions or associated images.
The text was updated successfully, but these errors were encountered: