AI Interview Platforms Tested: CodeSignal, Humanly, Eightfold in Job Screening

The Verge's senior AI reporter Hayden Field tested three AI interview platforms for job screening: CodeSignal, Humanly, and Eightfold. These platforms use AI avatars to conduct one-on-one video interviews with job applicants, asking questions and analyzing responses.
How AI Interview Platforms Work
The AI tools operate by having applicants participate in video calls where an AI avatar asks questions and evaluates responses. Companies behind these platforms claim they allow organizations to interview virtually every applicant for initial screening rather than just a subset. Some argue these systems analyze responses rather than visual cues, potentially reducing bias.
Limitations and Challenges
Despite claims of reduced bias, the article notes that bias-free AI systems are impossible to achieve. The models are trained on large internet datasets containing sexism, racism, and other biases. Field reported that while some platforms felt more natural than others, each time she wished she was talking to a human instead. She specifically mentioned struggling with the "uncanny valley" effect of looking at an AI avatar listening to her answers.
Testing Methodology
Field tested the platforms for various jobs, including positions created for the exercise based on her current role and real jobs listed at Vox Media. The testing revealed differences in how natural each platform felt, though all shared the fundamental limitation of being AI-driven rather than human-conducted interviews.
📖 Read the full source: HN AI Agents
👀 See Also

Georgia AI Data Center Drained 29M Gallons of Unmetered Water
QTS Fayetteville campus drew 29M gallons via two unauthorized water connections over 15 months, causing low pressure complaints. County waived fines, charged $147K retroactive.

VibeThinker-3B: A 3B Parameter Model That Matches 671B DeepSeek on AIME Math Benchmarks
Sina Weibo researchers released VibeThinker-3B, a 3B parameter model scoring 94.3 on AIME 2026—matching DeepSeek V3.2 (671B). The paper introduces the Parametric Compression-Coverage Hypothesis, arguing verifiable reasoning can be compressed into small models.

KV Cache Architecture Evolution: From GPT-2 to Mamba
Analysis of KV cache memory costs shows GPT-2 used 300 KiB/token, Llama 3 reduced it to 128 KiB/token with grouped-query attention, and DeepSeek V3 achieved 68.6 KiB/token with multi-head latent attention. Mamba/SSMs eliminate KV cache entirely with fixed-size hidden states.

CC 2.1.128 Release: New Built-in Background Agent, C# Beta Support, and Model Deprecations
CC 2.1.128 (+1406 tokens) adds built-in background-agent instructions, C# tool-runner/Managed Agents beta support, deprecates Sonnet 4 and Opus 4 recommending Opus 4.7/Sonnet 4.6, and removes session memory templates.