Welcome to RTC-SDD Challenge

The first speech deepfake detection challenge in real-time communication.

Recent advances in speech synthesis have made it increasingly easy to generate highly natural and identity-consistent speech. In online communication scenarios, where voice is often used for identity authentication and trust establishment, maliciously generated speech can be exploited for impersonation, fraud, and social engineering, posing serious threats to personal privacy, property security, and broader societal security.

Although substantial progress has been made in speech deepfake detection, existing research and benchmarks have mainly focused on offline audio or simulated transmission conditions. In real-time communication (RTC) scenarios, speech undergoes platform-specific enhancement, encoding, compression, network transmission, decoding, and reconstruction before reaching the receiver. These cascaded and coupled distortions can weaken or destroy the fine-grained artifacts relied upon by deepfake detection systems. Moreover, the internal speech processing pipelines of mainstream RTC platforms are typically black boxes, making such real-world distortions difficult to reproduce through simulation alone.

The RTC-SDD (Real-Time Communication Speech Deepfake Detection) Challenge aims to promote the development of robust and generalizable speech deepfake detection methods for realistic RTC environments. We welcome participants from academia and industry to join the challenge and work together to advance research in speech security.

As part of the ICASSP 2027 Signal Processing Grand Challenge program, the top five teams will be invited to submit a two-page paper and present their work in person at ICASSP 2027. Accepted papers will be published in the ICASSP proceedings and must be covered by an ICASSP registration. Teams presenting in person will also be invited to submit a full-length paper to the IEEE Open Journal of Signal Processing (OJ-SP), subject to peer review.

Challenge 1

Black-box Processing

Internal processing pipelines of RTC platforms are not exposed, making their behavior difficult to analyze.
Challenge 2

Robustness

Shifting from traditional noise robustness to robustness against distortions introduced by the speech processing pipeline in RTC transmission.
Challenge 3

Generalizability

Generalizing to unseen speakers, unseen generators, and unseen RTC platforms.

Challenge Timeline

Keep track of the key dates for challenge participation and submission.

September 01, 2026
Progress evaluation stage begins; progress evaluation data released.
November 09, 2026
Final evaluation stage begins; final evaluation data released.
November 16, 2026
Final submission deadline.
November 23, 2026
Challenge results and final ranking announced.
January 07, 2027
Two-page paper submission deadline for invited teams.
January 21, 2027
Two-page paper acceptance notification.
January 28, 2027
Camera-ready two-page paper submission deadline.

* The first four dates follow AoE; paper-related dates follow US Pacific Time, with deadlines at 11:59 PM.

Contact Us

For registration, dataset access, submission format, Codabench issues, or challenge rules, contact the organizing team.

rtcsddchallenge2027@gmail.com

Leaderboard

Track results and rankings during the challenge period.

Registered Teams

0 teams registered · Last updated: unknown

Loading teams...

Progress Evaluation Stage

Updated after the progress evaluation stage closes.

Rank Team Organization Clean Macro-F1 Noisy Macro-F1 Final Macro-F1
Not available yet

Final Evaluation Stage

Updated after the final evaluation stage closes.

Rank Team Organization Clean Macro-F1 Noisy Macro-F1 Final Macro-F1
Not available yet

RTC-SDD Challenge

Speech Deepfake Detection Challenge in Real-Time Communication

Why RTC-SDD?

Most existing speech deepfake detection benchmarks focus on offline audio files. However, real-world attacks often take place in online meetings and social communication scenarios. In these cases, speech is transformed by RTC pipelines before reaching the receiver, and detection systems must operate on the transmitted signal rather than the original generated waveform.

The RTC-SDD Challenge investigates this more realistic setting. It evaluates whether detection systems can remain reliable when generation artifacts are degraded, reshaped, or entangled with platform-specific communication artifacts.

Scenario

Post-RTC detection

Detection is performed on speech after RTC transmission, where spoofing cues are degraded by platform processing and network transmission.

Goal

Robustness and Generalizability

Build robust speech deepfake detection systems that generalize across generators, noise conditions, and RTC platforms.

RTC-SDD Dataset

RTC-SDD construction pipeline: speech construction, noise simulation, and real-time communication

Step 1: Speech Construction. Real speech is collected and spoofed speech is generated using ten state-of-the-art TTS/VC synthesis technologies, covering diverse generation mechanisms and spoofing artifacts.

Step 2: Noise Simulation. After dataset partitioning into training, development, and evaluation subsets, environmental noise simulation is applied to the evaluation set, including echo, keyboard, footsteps, rain, office, and coffee-shop sounds.

Step 3: RTC Transmission. The resulting speech data are transmitted through seven mainstream social media and meeting platforms, including QQ, Zoom, WeChat, DingTalk, Lark, VooV, and Telegram. This produces RTC-transmitted online samples that better reflect real-world communication scenarios.

Task Definition

Participants submit a spoof-likelihood score for each utterance ID in the official evaluation list, where higher scores indicate a higher probability of spoofed speech. Labels remain binary, bonafide or spoof, regardless of whether the audio is offline or RTC-transmitted.

The organizers provide training and development data with both offline speech before transmission and online speech after transmission. The evaluation set contains only online speech after RTC transmission. It includes a clean online subset and a noisy online subset; noisy speech is reserved for evaluation and is not included in the training or development sets. Some evaluation samples are transmitted through unseen RTC platforms.

Evaluation Metric

The official evaluation metric is Macro-F1 over the spoof and bonafide classes:

Macro-F1 = 1/2 (F1sp + F1bf).

The clean and noisy online subsets are evaluated separately, with their Macro-F1 scores denoted as Fclean and Fnoisy. The final score is computed as:

Macro-F1final = λ Fclean + (1 - λ) Fnoisy, where λ = 0.3.

This weighting assigns more importance to the noisy subset, encouraging robust detection under challenging RTC conditions.

Instructions

Recommended participation workflow from registration to final submission.

1

Register your team

  • If you do not already have a Codabench account, please create one first.
  • Complete the team registration form with your team name, members’ names, contact email addresses, affiliation, and Codabench username. All registration information must be provided in English.
  • After submitting the team registration form, please apply to participate in the challenge on Codabench. The organizers will review the registration information submitted through both the registration form and Codabench.
  • Teams whose registrations have been approved will be listed on the Leaderboard page of the official website. If you notice any incorrect information, experience a delay in the review process, or encounter any other issues, please contact the organizers.
Registration Form Codabench
2

Download dataset and baseline

Use the official RTCFake dataset and baseline code as the starting point. Keep train, development, progress, and evaluation partitions logically separated.

3

Train and validate your system

Train on the official training data, tune on the development set, and report validation performance. Avoid using progress or final evaluation data for tuning or analysis.

4

Submit predictions on Codabench

Prepare a two-column score file using the required format: utterance_id score. Scores should be normalized to [0, 1].

Codabench

Rules

Guidelines for participation, data use, submissions, and reproducibility.

1

Participation

Participants may join individually or as a team. Each participant can be affiliated with only one team during the challenge.

2

Data usage

The official training and development sets can be used for model training, validation, and model selection. The use of external speech datasets for training or fine-tuning is not permitted. Progress and final evaluation sets must not be used for training, parameter tuning, model selection, pseudo-labeling, manual annotation, or manual analysis.

3

Data augmentation

Participants may use public data augmentation methods or design their own augmentation strategies, but any external data introduced during augmentation must come from publicly available non-speech datasets. Official data must not be transmitted through RTC platforms for augmentation.

4

External resources

Publicly released model weights are permitted, provided that they have not been trained, fine-tuned, or post-trained on any external spoofing data. Closed-source model services, proprietary model APIs, and unpublished private weights are not allowed.

4

Model ensemble

Model ensembles are not permitted. Each submission must be generated by a single model and must not combine predictions, scores, or outputs from multiple independently trained models.

5

Submission policy

Submissions should follow the official Codabench format. Each score must correspond to exactly one utterance in the official evaluation list. Submissions should be provided as a two-column text file, with each line formatted as "id score".

6

Reproducibility

Top-ranked teams invited to submit papers should provide sufficient technical details, including model architecture, training strategy, augmentation, external resources, and hyperparameters.

FAQs

Answers to frequently asked questions

Registration Is there a deadline for registration?

There is no fixed registration deadline. Teams may register at any time while the competition is ongoing.

Registration Can I update my registration information after submission?

Yes. If you need to update your team name, team members, email address, or any other registration information, please submit a new registration form with the updated information.

Please specify the changes in the Note field. This applies regardless of whether your previous registration has already been approved.

The organizing committee will review the updated information as soon as possible. Once the update has been approved, you will receive a new confirmation email.

Registration When will my RTCFake dataset access request be approved?

Dataset access requests from successfully registered teams will normally be approved within 24 hours.

If your request has not been approved after 24 hours, please check that:

  1. You have received the registration confirmation email indicating that your competition registration has been successfully approved.
  2. The information provided in your Hugging Face access request is consistent with your competition registration information.

If both conditions are satisfied and your request is still pending, please contact the organizing committee at rtcsddchallenge2027@gmail.com.

Rule Clarification Can publicly released model weights trained, fine-tuned, or post-trained on external spoofing data be used?

No. Models that have been trained, fine-tuned, or post-trained on external spoofing data are not permitted, even if the weights are publicly available.

For example, DF_Arena_1B_V_1 is not eligible for use in this competition because it has been trained on external spoofing data.

Rule Clarification What is considered an RTC platform, and is local simulation of RTC-related processing allowed?

In the competition rules, "RTC platforms" refer to real-world real-time communication platforms or services, such as Zoom, WeChat, and QQ. The official data must not be transmitted through these platforms for data augmentation.

Local and offline simulation of RTC-related processing is allowed, provided that the official data is not actually transmitted through a real-world RTC platform and that the augmentation otherwise complies with the competition rules.

For example, participants may use open-source WebRTC audio processing modules, such as AEC, NS, and AGC, or other locally implemented algorithms or processing pipelines to simulate RTC-related effects.

Evaluation How is the threshold determined when calculating Macro-F1?

A fixed threshold of 0.5 is used when calculating Macro-F1.

Rule Clarification Can the speech subset of MUSAN be used for data augmentation?

No. According to the competition rules, any external data used for augmentation must come from publicly available non-speech datasets. Therefore, the speech subset of MUSAN is not permitted.

Rule Clarification Can publicly available models be used for data augmentation?

Yes. Publicly available models or tools may be used to process the official data for augmentation, provided that their use complies with the external resource rules.

Registration Are all team members required to have the same affiliation, and how should affiliation information be provided?

Team members are not required to have the same affiliation. In the registration form, you may provide only the primary affiliation or include the affiliations of all team members. If an individual has multiple affiliations, providing one of them in the registration form is sufficient.

For subsequent papers, author affiliations should follow the actual authorship information and the applicable affiliation requirements.

Rule Clarification Is weight averaging across different training runs of the same model allowed?

Yes. Weight averaging across different training runs of the same model is permitted, provided that the final submission results in and uses a single model for inference. Combining predictions, scores, or outputs from multiple models remains prohibited.

Rule Clarification Can official bonafide audio be resynthesized to generate additional spoof training data?

No. Using vocoders, voice-conversion models, or similar methods to generate additional spoof speech samples from the official data is not permitted.

Evaluation Will the evaluation protocol disclose condition or language information for individual utterances?

No. The evaluation protocol will not disclose condition or language information for individual utterances.

Evaluation Will the evaluation set contain offline audio?

No. The evaluation set will contain only transmitted online audio and will not include offline audio.

Evaluation How can I view the Clean Macro-F1 and Noisy Macro-F1 scores for an individual submission?

You can view the detailed scores of an individual finished submission through the Codabench submission API:

https://www.codabench.org/api/submissions/<Submission ID>/

While logged in to Codabench, replace <Submission ID> with the corresponding submission ID. The returned scores field includes the Clean Macro-F1 and Noisy Macro-F1 scores for that submission.

Organization

Meet the organizing team of the RTC-SDD Challenge.

Zhen-hua Ling
Zhen-Hua Ling
University of Science and Technology of China, China
Professor

Questions about RTC-SDD?

For any question about registration, dataset access, baseline code, submission format, or rules, contact the RTC-SDD Challenge team.

rtcsddchallenge2027@gmail.com