<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"><channel><title>Mikita Daroshkin</title><description>Mikita Daroshkin, senior AI/ML engineer, forward-deployed. Agentic AI that Fortune 500 companies run in production under audit: pharma, healthcare, lending.</description><link>https://mikitadaroshkin.com/</link><item><title>A confidence score is not a reason</title><link>https://mikitadaroshkin.com/posts/a-confidence-score-is-not-a-reason/</link><guid isPermaLink="true">https://mikitadaroshkin.com/posts/a-confidence-score-is-not-a-reason/</guid><description>An extraction confidence is not an adverse-action reason. Document Intelligence returns two scores, and the gate reads the wrong one.</description><pubDate>Tue, 21 Jul 2026 16:00:00 GMT</pubDate></item><item><title>The permission model was built for people holding jobs</title><link>https://mikitadaroshkin.com/posts/agents-inside-a-system-of-record/</link><guid isPermaLink="true">https://mikitadaroshkin.com/posts/agents-inside-a-system-of-record/</guid><description>An HCM grants permissions to people holding jobs, and an agent is neither. Least privilege with no native expression, and the primitive that just shipped.</description><pubDate>Tue, 07 Jul 2026 16:00:00 GMT</pubDate></item><item><title>Nobody owns the join between two agent platforms</title><link>https://mikitadaroshkin.com/posts/federating-agents-across-vendor-platforms/</link><guid isPermaLink="true">https://mikitadaroshkin.com/posts/federating-agents-across-vendor-platforms/</guid><description>Four vendor agent platforms in one mesh, each sound alone. The seven joins that failed, starting with three catalogs where the loosest one decides.</description><pubDate>Tue, 23 Jun 2026 16:00:00 GMT</pubDate></item><item><title>The LLM judge is a model too</title><link>https://mikitadaroshkin.com/posts/llm-judge-is-a-model-too/</link><guid isPermaLink="true">https://mikitadaroshkin.com/posts/llm-judge-is-a-model-too/</guid><description>The judge is a model too, with the same length, position and self-preference biases. Measuring it, versioning it, and knowing when not to use one.</description><pubDate>Tue, 16 Jun 2026 16:00:00 GMT</pubDate></item><item><title>Containment is measured on the calls that never needed you</title><link>https://mikitadaroshkin.com/posts/when-the-agent-spends-the-customers-money/</link><guid isPermaLink="true">https://mikitadaroshkin.com/posts/when-the-agent-spends-the-customers-money/</guid><description>Containment counts the sessions that never needed you. A booking agent&apos;s rollback is a refund, and a surge multiple is a scalar over a rotating mixture.</description><pubDate>Tue, 09 Jun 2026 16:00:00 GMT</pubDate></item><item><title>Throughput per dollar, once the traffic is bursty</title><link>https://mikitadaroshkin.com/posts/throughput-per-dollar-serving-under-bursty-traffic/</link><guid isPermaLink="true">https://mikitadaroshkin.com/posts/throughput-per-dollar-serving-under-bursty-traffic/</guid><description>One roofline number sets the speculative-decoding gate, the batch-size knee and the chunked-prefill floor. Why max-num-seqs is a latency control.</description><pubDate>Tue, 26 May 2026 16:00:00 GMT</pubDate></item><item><title>Validating a system that never returns the same answer twice</title><link>https://mikitadaroshkin.com/posts/validation-as-code-under-gxp/</link><guid isPermaLink="true">https://mikitadaroshkin.com/posts/validation-as-code-under-gxp/</guid><description>Validating a system that never returns the same answer twice: what qualification becomes when the output is distributional, and where sampling breaks.</description><pubDate>Tue, 05 May 2026 16:00:00 GMT</pubDate></item><item><title>The guarantee is marginal and the audit is not</title><link>https://mikitadaroshkin.com/posts/calibrated-risk-when-a-person-is-the-outcome/</link><guid isPermaLink="true">https://mikitadaroshkin.com/posts/calibrated-risk-when-a-person-is-the-outcome/</guid><description>A conformal coverage guarantee is marginal, not conditional, so it can hold overall and fail on exactly the subgroup your fairness audit is about.</description><pubDate>Tue, 14 Apr 2026 16:00:00 GMT</pubDate></item><item><title>The wall an 8B model hits on a mid-tier Android phone</title><link>https://mikitadaroshkin.com/posts/shipping-an-8b-model-offline-on-mid-tier-android/</link><guid isPermaLink="true">https://mikitadaroshkin.com/posts/shipping-an-8b-model-offline-on-mid-tier-android/</guid><description>An 8B model on a $250 Android phone hits a bandwidth ceiling, a residency ceiling and a thermal ceiling, in that order, and dies mid-sentence.</description><pubDate>Tue, 24 Mar 2026 16:00:00 GMT</pubDate></item><item><title>Tail sampling has three ways to lose the trace you wanted</title><link>https://mikitadaroshkin.com/posts/opentelemetry-span-design-for-agents/</link><guid isPermaLink="true">https://mikitadaroshkin.com/posts/opentelemetry-span-design-for-agents/</guid><description>Tail sampling fixes head sampling and adds three quieter ways to lose the same traces, none visible in the config and none logged as an error.</description><pubDate>Tue, 17 Feb 2026 16:00:00 GMT</pubDate></item><item><title>Model Armor and DLP have to get through the perimeter too</title><link>https://mikitadaroshkin.com/posts/agent-tools-inside-an-egress-locked-perimeter/</link><guid isPermaLink="true">https://mikitadaroshkin.com/posts/agent-tools-inside-an-egress-locked-perimeter/</guid><description>The content screens you put inline are themselves services behind the same wall. The endpoint Model Armor needs, and the field that says a screen ran.</description><pubDate>Tue, 13 Jan 2026 16:00:00 GMT</pubDate></item><item><title>Who runs the OAuth flow decides your network topology</title><link>https://mikitadaroshkin.com/posts/identity-delegation-across-agent-hops/</link><guid isPermaLink="true">https://mikitadaroshkin.com/posts/identity-delegation-across-agent-hops/</guid><description>Who runs the OAuth flow decides your network topology. Token exchange across agent hops, and why an act claim records who acted but never when.</description><pubDate>Tue, 11 Nov 2025 16:00:00 GMT</pubDate></item><item><title>The handoff is the part nobody writes down</title><link>https://mikitadaroshkin.com/posts/the-discipline-of-leaving/</link><guid isPermaLink="true">https://mikitadaroshkin.com/posts/the-discipline-of-leaving/</guid><description>Leaving a fleet of LLM agents a client&apos;s own team can run: discovery by directory, scorecards a stranger can read, runbooks drilled against real failures.</description><pubDate>Tue, 09 Sep 2025 16:00:00 GMT</pubDate></item><item><title>Faithful to the wrong version</title><link>https://mikitadaroshkin.com/posts/rag-on-a-regulated-corpus/</link><guid isPermaLink="true">https://mikitadaroshkin.com/posts/rag-on-a-regulated-corpus/</guid><description>On a versioned corpus, faithfulness metrics pass an answer grounded in a revision retired months ago. Treating the corpus as bitemporal is what fixes it.</description><pubDate>Tue, 17 Jun 2025 16:00:00 GMT</pubDate></item><item><title>Text-to-SQL that can&apos;t leak data it shouldn&apos;t</title><link>https://mikitadaroshkin.com/posts/text-to-sql-that-cant-leak/</link><guid isPermaLink="true">https://mikitadaroshkin.com/posts/text-to-sql-that-cant-leak/</guid><description>The seam between what a model may see and what it may run, enforced by different systems owned by different teams, and the error message as a leak channel.</description><pubDate>Tue, 18 Mar 2025 16:00:00 GMT</pubDate></item></channel></rss>