On the instructs2s-eval benchmark in the offline scenario, SpeechGPT achieves an ASR-WER of 45.00, far higher than all other systems, indicating that its generated speech is poorly aligned with its own text response. Its ChatGPT score is 2.98 for speech-to-text and drops to 2.19 for speech-to-speech, the lowest among all evaluated models. In the streaming scenario, SpeechGPT's decoding latency exceeds 4500 ms due to sequential generation of text instruction, text response, and speech units, and its quality metrics remain the worst across all latency settings.