We’re excited to release
@basiscompany’s first open source dataset: 1,500 hrs of track separated, multi-speaker conversations across 22 languages. The largest ever of its kind and first for several languages.
Basis has built up a network of over 1 million contributors submitting audio, video, annotations, ratings, transcriptions and will keep open-sourcing interaction data to cover all the weird/ambiguous/natural modes of human interaction (more soon!).
This release includes 70+ of hours with 3+ live speakers and 2645 native speakers meeting, sharing, arguing, laughing, trolling in English, Spanish, Japanese, Korean, Georgian, Mingrelian, German, French, Italian, Hebrew, Russian, Hindi, Arabic, Ukrainian, Dutch, Xhosa, Zulu, Portuguese, Chinese, Turkish, Polish, and Kazakh.
For some of these languages, this is the first ever open release of duplex conversational data.
100 hrs are densely labeled with 113k human judgements on subtleties that models struggle with (was this backchannel sympathetic or frustrated? Was this silence awkward or turn-holding? Was this utterance directed or general? …)
Reach out if you’re interested in more!
ჯგირი მორაგადეს ჯგირი მარჩქილე ოკონია