LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing

EACL 2026 · Stanford Autonomous Agents Lab · 26 citations

Daniel Fein, Sebastian Russo, Violet Xiang, Kabir Jolly, Rafael Rafailov, Nick Haber

Do machines have the same quality preferences for creative writing as humans? LitBench is a benchmark for finding out: 2,480 human-labeled story comparisons plus a training corpus of ~43,000 preference pairs. We benchmarked LLMs as zero-shot judges of creative writing and trained dedicated reward models — the trained models reach 78% agreement with human preferences, beating the best off-the-shelf judge (73%), and we validated the results with human studies on newly generated stories.

Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini

Google DeepMind · May 2026

Gemini Embedding 2 embeds video, audio, image, and text into a single shared representation space, so arbitrary interleaved combinations of those modalities can be compared directly. Trained with large-scale contrastive learning in a multi-task, multi-stage setup, it reaches state-of-the-art results on unimodal, cross-modal, and multimodal retrieval benchmarks.