RUT-Bench Benchmark data in "Beyond Ideal Instruction: A Comprehensive Framework for Evaluating LLMs in Realistic Interactions". Miaow-Lab/RUT-Bench Viewer • Updated Jun 4 • 1.64k • 141 • 1 Beyond Ideal Instruction: A Comprehensive Framework for Evaluating LLMs in Realistic Interactions Paper • 2606.03318 • Published Jun 2
Beyond Ideal Instruction: A Comprehensive Framework for Evaluating LLMs in Realistic Interactions Paper • 2606.03318 • Published Jun 2
SSAE Training and evaluation dataset, model checkpoints in 'Step-Level Sparse Autoencoder for Reasoning Process Interpretation' Miaow-Lab/SSAE-Dataset Viewer • Updated Mar 4 • 1.28M • 62 Miaow-Lab/SSAE-Checkpoints Feature Extraction • Updated Mar 4 Step-Level Sparse Autoencoder for Reasoning Process Interpretation Paper • 2603.03031 • Published Mar 3
Step-Level Sparse Autoencoder for Reasoning Process Interpretation Paper • 2603.03031 • Published Mar 3
RUT-Bench Benchmark data in "Beyond Ideal Instruction: A Comprehensive Framework for Evaluating LLMs in Realistic Interactions". Miaow-Lab/RUT-Bench Viewer • Updated Jun 4 • 1.64k • 141 • 1 Beyond Ideal Instruction: A Comprehensive Framework for Evaluating LLMs in Realistic Interactions Paper • 2606.03318 • Published Jun 2
Beyond Ideal Instruction: A Comprehensive Framework for Evaluating LLMs in Realistic Interactions Paper • 2606.03318 • Published Jun 2
SSAE Training and evaluation dataset, model checkpoints in 'Step-Level Sparse Autoencoder for Reasoning Process Interpretation' Miaow-Lab/SSAE-Dataset Viewer • Updated Mar 4 • 1.28M • 62 Miaow-Lab/SSAE-Checkpoints Feature Extraction • Updated Mar 4 Step-Level Sparse Autoencoder for Reasoning Process Interpretation Paper • 2603.03031 • Published Mar 3
Step-Level Sparse Autoencoder for Reasoning Process Interpretation Paper • 2603.03031 • Published Mar 3