Sitemap
Main Pages
Browse by Tag
All Posts
- OpenAgentFlow: Enabling System-Wide Safety Boundaries for Heterogeneous AI Agent Fleets
- Incremental Risk Assessment of Progressive Elder Financial Scams via Instruction-Tuned Small Language Models
- UI-Venus-2 Technical Report
- I-CARE: Analysis of interference-related phenomena in a controllable, diverse and representative unlearning setting for text-to-image models
- HyperWorld: Hypergraph-Structured State Serialization Improves Learned Textual World Models
- When Prediction Error Is Not Enough: Evaluating Nuisance-Function Prediction for Causal Estimation
- Discrete-Time MDP Modeling for Multi-Item Capacitated Lot Sizing with Stochastic Demand Timing
- Long-Horizon State Tracking in LLMs: Executing MD5 through a Deep Sequence of Dependent Tool Calls
- SCAFFOLD: A Large-Scale Structured Dataset of Computer Science Research Figures with Diagram QA and Chain-of-Thought Reasoning Traces
- EULER: Exploring Underused Links with Evidence-Checked Return for Multi-Agent Mathematical Discovery
- Expert-validated STEM QA
- Paper Pilot: A Human-in-the-Loop Expert System for Evidence-Traceable Scientific Manuscript Generation in Applied Sciences
- Introducing @huggingface/kernels: 200+ WebGPU Kernels for Local AI
- BenchMIRT: What are LLM benchmarks actually measuring?
- Statutory AI: Aligning Large Language Models With Legal Norms
- The Signal in the Noise: An Auditable Reliability Layer for Biomedical Text Classification
- From Question-First to Analyst-First: Domain-Expert Skills and Verified Knowledge Compilation for Proactive Enterprise Analytics
- The Race between Agentic AI Capabilities and Data Quality Control in Online Surveys
- DS-Lighting: Making Agent Harnesses Explicit for Data-Science Automation
- CDPR: Counterfactual Advantage-based Credit Assignment for Cost-Aware Sequential Medical Diagnosis
- SHAPE of Chain-of-Thought in Math Reasoning
- A collective capability boundary in frontier large language models on guideline-conformant and case-specific oncology decision-making
- The Open ASR Leaderboard Adds Its First Global South Language
- Training and Finetuning Multi-Vector Embedding Models with Sentence Transformers
- Wire It, Run It, Deploy It: AI Workflows in Gradio
- Granite 4.2 LLMs: How They're Built
- Quantization-Aware Healing: a compressed, 4-bit model that outperforms its full-precision original
- Measuring benchmark optimization in speech recognition
- How Hugging Face Inference Endpoints, Jobs, and Buckets Power Search on Papers with Code
- Up to 3.2x Faster Inference with LFM2.5-DSpark
- Welcome to Your Blog