Security
Anthropic improves alignment and security efforts
Resumed external cyber evals, a real-time escape and probe classifier, and partner sandbox guidance. Affects how Claude agents are scoped and monitored.
AI news
Security
Anthropic improves alignment and security efforts
Resumed external cyber evals, a real-time escape and probe classifier, and partner sandbox guidance. Affects how Claude agents are scoped and monitored.
Models
DeepSeek publishes V4-Flash-Vision-Exp weights
MIT-licensed ~305B multimodal MoE weights with vLLM and SGLang serve notes for screenshot, chart and document workflows.
Agents
Google Antigravity Teamwork posts multi-agent wins
Multi-agent Teamwork updates claimed open theory results verified in Lean, a RISC-V simulator that boots xv6, and upstream Eigen and ParlayHash optimisations.
Language
Salesforce open-sources ClaimProbe and ClaimWriter
Claim-level audit cut hallucination about 2.6 to 4.5 times in deep-research hosts where rubric scores can hide claim failures.
Language
Sony and UCLA post AToM CoWriter writing support
Infers writing processes from keystrokes and fires support agents so writers need not craft prompts.