A 0.9B model for long-form transcription in 50+ languages with speaker diarization, timestamps, and acoustic event awareness