<?xml version="1.0" encoding="utf-8"?>
<feed xmlns="http://www.w3.org/2005/Atom">
  <generator uri="https://jekyllrb.com/" version="4.4.1">Jekyll</generator>
  <link href="https://yongzx.github.io/feed.xml" rel="self" type="application/atom+xml" />
  <link href="https://yongzx.github.io/blog/" rel="alternate" type="text/html" />
  <updated>2026-07-29T23:48:46-07:00</updated>
  <id>https://yongzx.github.io/feed.xml</id>
  <title type="html">Yong Zheng-Xin</title>
  <subtitle>Yong Zheng-Xin&apos;s personal website</subtitle>
  <author>
    <name>Yong Zheng-Xin</name>
    <email>contact.yong@brown.edu</email>
    <uri>https://yongzx.github.io</uri>
  </author>

  
  
  
  
    
      
      
      
        
        
        <entry>
          <title type="html">Sam Altman on AGI, Compute, and Human Agency</title>
          <link href="https://yongzx.github.io/blog/2026/07/29/sam-altman-on-agi-compute-and-human-agency/?utm_source=rss&amp;utm_medium=feed&amp;utm_campaign=blog" rel="alternate" type="text/html" title="Sam Altman on AGI, Compute, and Human Agency" />
          <published>2026-07-29T00:00:00-07:00</published>
          <updated>2026-07-29T00:00:00-07:00</updated>
          <id>https://yongzx.github.io/blog/2026/07/29/sam-altman-on-agi-compute-and-human-agency/</id>
          
            <content type="html" xml:base="https://yongzx.github.io/blog/2026/07/29/sam-altman-on-agi-compute-and-human-agency/"><![CDATA[<p>Disclaimers:</p>
<ul>
  <li>I haven’t started at OpenAI as this post is written.</li>
  <li>Tenses are used fluidly (present tense to indicate what I think he believes in right now).</li>
  <li>No AI-assisted writing is used (for better or worse 🤷‍♂️).</li>
  <li>
    <bs>Blue texts</bs>
    <p>are my thoughts or interpretations.</p>
  </li>
</ul>

<hr />

<h3 id="the-past-year-was-indeed-very-tough-and-thats-partly-my-responsibility-but-the-year-ahead-could-be-our-best-twelve-months">“The past year was indeed very tough, and that’s partly my responsibility. But the year ahead could be our best twelve months”</h3>

<p><strong><redspan>What went wrong:</redspan></strong> Sam said that they spreaded themselves too thin. Last year, they were uncertain whether revenue would grow quickly enough to justify the large compute commitments, so they explored consumer applications and other ways to monetize underused GPUs. Once model progress and economic demand became clearer, the company narrowed the focus to delivering the best intelligence at the lowest possible cost and enabling others to build on top of it.</p>

<p><strong><redspan>What's coming next:</redspan></strong> Better models and better products built around them. They will continue to focus on (1) training great models that can be used in different ways that can generate economic value; (2) producing or partnering to manufacture the chips and systems, and finding a place to house them; and (3) building robots that can automate this process to continuously drive down costs across the entire supply chain.</p>

<blockquote>
  <p>Fundamentally, our business is about selling AI that enables people to build extraordinary products and services for each other using these components.  … Building every vertical application ourselves, trying to compete with every venture and every company? We have zero interest in that. We genuinely just want to provide the platform. –– Sam Altman</p>
</blockquote>

<h3 id="convictions-and-worries">Convictions and Worries</h3>
<p><strong><redspan>Conviction on securing compute and building datacenters:</redspan></strong> Sam expected the exponential model improvement especially after seeing GPT-4, where he believed that reasoning once solved will lead to agents that do huge economic work. He also expected huge demands and regardless of how efficient the models would become, compute would remain critical. That led to OpenAI trying to talk to cloud providers, chip manufacturers, and energy suppliers to secure unprecedentedly huge compute capacity (which many touted at that time as insane and reckless). They only got two yes––from Microsoft and Oracle––but that allowed them to move forward.</p>

<p>Interesting thing I’ve learned here is that data centers can now be built in deserts. This is possible because modern data centers use closed-loop water system and only as much water as an office building. They can also be powered by nuclear or solar energy instead of fossil fuels. <bs>This totally surprised me because most environmental concerns against building data centers are addressed.</bs></p>

<p>OpenAI’s goal is to serve users with the best intelligence-price trade-offs, and Sam seemed unperturbed by the <a href="https://www.anthropic.com/news/detecting-and-preventing-distillation-attacks">distillation attacks</a>. To Sam, at sufficient scale, inference revenue (even at modest margins) can cover the cost of training frontier models.</p>

<p><strong><redspan>His top worry is about AI safety:</redspan></strong> He was shooked by the <a href="https://huggingface.co/blog/agent-intrusion-technical-timeline">Huggingface incident</a>, where the OpenAI’s internal models hacked out of the sandbox by chaining together multiple zero-day exploits and breached Huggingface platforms to find answers for ExploitGym evals. As a result, the team had paused training and began working on stronger sandbox security; in the long-term, he thinks that we’d need to pace AI development to buy time for society to harden infrastructure around these unintended behaviors from capability progress. He is also aware of how these might come across as regulatory capature or collusion among frontier labs, so he urges for solutions for this.</p>

<p>Later on, he’s also worried about concentration of power––particularly terrified by a world where only a small group of people have access to AI and make decisions around it. He warned about not falling into the trap of AI safety and understandable fears, and growing up where internet has no rules, he wants to preserve that spirit with AI technology and let people self-determine the future.</p>

<h3 id="mission-and-bottlenecks">Mission and Bottlenecks</h3>

<p><strong><redspan>OpenAI's mission and "AGI being a genie that can grant wishes"</redspan></strong>: Sam thinks that AI will become a genie that can grant wishes, and he hopes that OpenAI will continue to</p>
<ol>
  <li>Get them into everyone’s hands and make people’s lives better than they otherwise they have been.</li>
  <li>Make sure people maintain control and agency.</li>
</ol>

<p>He thinks that it will lead to more jobs and help more creative endeavors.</p>

<p>However, he doesn’t think that AGI is here yet because right now it cannot do tall orders such as “curing cancer” and more importantly, it suffers from limitations such as not being able to continuously learn. However, he sympathized with people who claim that AGI is already here as in this worldview, we already have a machinery that can produce better AI models where we are learning new science about them. In my opinion, I think this is not very well elaborated by Sam, as I think <bs>the machinery part is about the *speed* at which we are producing a much-better models.</bs></p>

<p>When asked about what people in 2019 would think about GPT 5.6, he said that they would agree it is AGI, but they would be overestimating the economic impacts especially with respect to AI taking jobs. He gave several potential reasons: AI is jagged so still complementary to humans, and humans enjoy interacting with humans. Citing examples about AI-generated arts, he believes that there’s value to person behind the creation process; and in business, that means we would prefer somebody that can be held accountable for and not an AI CEO.</p>

<p><strong><redspan>Bottleneck at frontier labs: data, compute, ideas, or talent?</redspan></strong> He thinks that it has been a cycle among research ideas, compute, and data. While compute is the primary constraint now, major breakthroughs over the past six months are due to research ideas. As RSI is approaching, he expects that the workflow of the researchers will change drastically like how software engineers don’t write code any more in the traditional sense, but it’s still important to have researchers to tell computers what they want. <bs>This indexes towards the importance of developing research taste and good intutitions.</bs></p>

<hr />

<h3 id="less-organized-notes">Less-Organized Notes</h3>

<blockquote>
  <p>Q: What’s it like becoming a dad and having growing kids in this era? <br />
A: I think I have the best, most interesting job in the world, and it is still a very distant second to having kids.</p>
</blockquote>

<p>I find the second half of the interview becomes really hard to summarized and grouped as above. It becomes more conversational. So here I try to use bullet points to capture interesting takes:</p>

<ul>
  <li>
    <p><strong>On being early:</strong> Best investment is made when it is not popular or following what other people are already doing. <bs>Honestly, I find this a bit hard to generalize because looking at AI progress, in my opinion you definitely should work on LLMs, which is a trendy topic. However, you should come in with your personal angle of attack where the reward-to-effort (or impact-to-effort) ratio is high.</bs></p>
  </li>
  <li>
    <p><strong>Adaptability of humans.</strong> He thinks people adapt very quickly, citing how in pandemic, lockdowns happen quickly and people adapt accordingly. <bs>This probably suggests that when singularity happens, even if chaos ensue, it will not last long and people can still adapt and flourish just like during pandemic.</bs></p>
  </li>
  <li>
    <p><strong>Evals that matters.</strong> Evals that matters is whether AI is useful to people. Right now, people approximate this via revenue, GDP, etc., and OpenAI has teams figuring out how to evaluate on superintelligent models.</p>
  </li>
  <li>
    <p><strong>Personal-AI.</strong> His vision of how future AI-human interactions may look like is that AI oversees everything the user does and continues to think when users are sleeping so once users wake up, their AI agent would present them new ideas or to-dos. This needs a lot of compute, and he believes that everyone like him would be willing to pay a lot for that always-on AI.</p>
  </li>
  <li>
    <p><strong>Taste.</strong> He believes it is still hard for future models to develop good taste.</p>
  </li>
  <li>
    <p><strong>Robotics.</strong> He thinks that ChatGPT moments for robotics would happen in like next two to three years, and that moment would look like “you type a command and a robot can do something crazy, and you can like watch it even if you are not physically there.”</p>
  </li>
  <li>
    <p><strong>ChatGPT release.</strong> Interestingly, OpenAI already had GPT-4 ready when they released ChatGPT, which they used a weaker version of the model (GPT-3.5) to avoid safety risks such as misuse. <br />Note that ChatGPT was released Nov 2022 and GPT-4 was released on March 2023.&lt;/br&gt; The whole idea of developing ChatGPT was from seeing how users organically like to interact with AI in a chatbot fashion, and they were amazed when they built that interface and interacted with GPT-4 internally.</p>
  </li>
</ul>

<blockquote>
  <p>We decided to build a proper chatbot. We started working on it, completed GPT-4, and began using it internally. We realized, ‘This is huge. This will be a real inflection point in how the world perceives AI.’ But there were thorny issues: ‘Will this spread disinformation? Will it say deeply offensive things? Will we get into trouble?’ So we opted to start with a weaker version. Releasing both the chat interface and GPT-4 simultaneously felt like too big a leap. Instead, we launched the chat interface powered by GPT-3.5.</p>
</blockquote>

<ul>
  <li>
    <p><strong>Marketing</strong>: Sam believes that great product markets itself simply based on utility. For user growth and value, smarter models, more compute, and better products will do. However, he thinks there’s a place for marketing: on where things are heading because people are feeling anxious about the technology.</p>
  </li>
  <li>
    <p><strong>Investment:</strong> <bs>I am personally surprised to learn that there's only one relentless all-in investor (who is always there) for OpenAI</bs> which is Josh Kushner.</p>
  </li>
</ul>

<blockquote>
  <p>Q: What have you learned about investors? <br />
A: The number of investors who will actually show up and help you is unbelievably small.</p>
</blockquote>

<ul>
  <li>
    <p><strong>Emotional state:</strong> When asked “what you don’t want me to know about you”, Sam thought for a few seconds and said that he was tired and repeat that it’s tiring. He didn’t reveal much on what keeps him going and commented that he would continue doing this for the rest of his career, but it’s really hard to explain to other people. <bs>Seeing how Lilian Weng quitting TML and joining OpenAI because she cannot continue at the pace a startup requires and wants a scoped role in a more predictable place, I have a lot more respect for Sam. While we can argue whether OpenAI should exist for existential AI safety reasons, it's undeniable that several scientific breakthroughs on AI come from early OpenAI era, and it's unimaginable how tough that is to navigate the field when giants like Yann did not believe in AGI.</bs></p>
  </li>
  <li>
    <p><strong>What would happen the month after ASI is achieved:</strong> Sam expects little immediate change. He rejects the belief in a “machine god” that would transform everything almost instantly. He argues that the pace of human progress has been consistent, which is a <em>smooth</em> exponential curve: enormous when zoomed out and viewed across decades, but experienced as a series of incremental steps.</p>
  </li>
</ul>

<blockquote>
  <p>Q: In the story of this company, who is your favorite unsung hero? <br />
A: The first person who comes to mind is Alec Radford. .. His work truly laid the foundation for the GPT series. Beyond many other critical contributions, he consistently inspired, guided, and pushed people toward directions that later proved immensely significant. … ‘He’s one of the kindest, most positive, and best people I’ve ever met.’</p>
</blockquote>

<ul>
  <li>
    <p><strong>Codex and competitive advantage</strong>: He thinks ChatGPT didn’t play a huge role in the Codex takeoffs; it’s just Codex being the best product with the best model. He believes in product advantage in such AI competition: if somebody builds something better than Codex, he foresees users could just as easily migrate away from Codex. He believes that intelligence would become a commodity, but the scale of compute clusters and the ability to produce more computing power are durable advantage.</p>
  </li>
  <li>
    <p><strong>New hardware</strong>: He wants to create something that feels natural to be in constant presence in a social setting (such as in the conversation he is having). The reason is as mentioned before, he wants an AI that can stay always-on, proactive, and understand all the user’s context.</p>
  </li>
  <li>
    <p><strong>Potential oversupply of AI compute</strong>: In a future in which highly efficient genie can satisfy all demand with little compute, limited human attention could become the bottleneck, leaving more compute available than people need.</p>
  </li>
  <li>
    <p><strong>Scaling laws</strong>: Lots of people want/predict it to break down, but it still keeps going.</p>
  </li>
  <li>
    <p><strong>Mistake of messing up org structure</strong>: Sam shared that their original innovation of org structure as non-profit had created substantial pain, even though back then it felt like the only viable move to protect OpenAI’s mission. <bs>I don't think he shared much about the reason why it is a mistake, and I couldn't find relevant resources online. I find his sharing especially relevant in light of <a href="https://www.lesswrong.com/posts/8T4MTntw4tyqc6eFm/don-t-default-to-nonprofit-1">AI safety grantfunders talking about how not to default to non-profit.</a></bs></p>
  </li>
</ul>

<blockquote>
  <p>Q: I always ask everyone the same traditional closing question: What’s the kindest thing someone has ever done for you? <br />
A: I feel incredibly fortunate that so many people have gone to such great lengths to be kind to me throughout my life. When reflecting on this, what comes to mind are all those moments––scattered across different parts of my life––where people showed me extraordinary kindness. Just yesterday, my child shared his blueberries with me for the first time. It was very sweet—a beautiful moment.</p>
</blockquote>]]></content>
          
          <author>
            <name>Yong Zheng-Xin</name>
          </author>
          
            <category term="notes" />
          
          
          
            <summary type="html"><![CDATA[Summary of Patrick O’Shaughnessy’s podcast with Sam Altman.]]></summary>
          
        </entry>
        
      
    
  
    
      
      
      
        
        
        <entry>
          <title type="html">Surprising lessons from my research scientist job search</title>
          <link href="https://yongzx.github.io/blog/2026/06/24/job-search/?utm_source=rss&amp;utm_medium=feed&amp;utm_campaign=blog" rel="alternate" type="text/html" title="Surprising lessons from my research scientist job search" />
          <published>2026-06-24T00:00:00-07:00</published>
          <updated>2026-06-24T00:00:00-07:00</updated>
          <id>https://yongzx.github.io/blog/2026/06/24/job-search/</id>
          
            <content type="html" xml:base="https://yongzx.github.io/blog/2026/06/24/job-search/"><![CDATA[<p>There are two recent blog posts from <a href="https://alisawuffles.github.io/blog/job-search/">Alisa</a> and <a href="https://silviasapora.github.io/blog/ml-interviews.html">Silvia</a>, both CS PhD students, on how they prepared and got into frontier labs such as OpenAI and Google Deepmind. I highly recommend them, and after seeing the <a href="https://x.com/ewveggies/status/2069211279366185027?s=20">reactions on Twitter</a>, I want to share a different angle: what surprised me during my own research scientist job search.</p>

<p>I write this post for two primary audiences:</p>
<ol>
  <li>CS PhD graduates who are probably like me, who spent 5-6 years working on multiple research papers and now trying to look for industry opportunities.</li>
  <li>AI safety fellows who are applying for full-time positions.</li>
</ol>

<p>Disclaimer: no LLMs were used in the writing.</p>

<h3 id="personal-experience">Personal Experience</h3>

<p>I am a fifth-year PhD student at Brown University. My job search experience is a bit unconventional as I did some research pivot during my last year of PhD.</p>

<p>In Fall 2025, I was applying for multilingual and AI safety positions, but mostly receiving research scientist opportunities for multilingual/post-training. This is because of my research portfolio having less work on core AI safety topics.</p>

<p>I decided during the semester that I needed to fully pivot to AI safety research because I think there are a lot of important areas within AI safety that need immediate attention as we are approaching AGI/ASI. So when I received the <a href="https://constellation.org/programs/astra">Astra Fellowship</a>, I decided to take a few month break from job search and focused on doing the fellowship well such that I am more qualified for higher-impact AI safety roles. To this end, I turned down existing offers and pushed back my graduation to 2027.</p>

<p>Nearing the end of my fellowship, I restarted my job search and things went a bit more disorganized than what I originally had in mind. My original plan was to finish my fellowship by June, turn my work into a paper, and start my interviews (which means I would have only started my interviews in July). However, due to timing reasons <sup id="fnref:timing"><a href="#fn:timing" class="footnote" rel="footnote" role="doc-noteref">1</a></sup> and my worry about headcount, I started in around mid-May and got offers that I am excited about before mid-June. I actually withdrew from some ongoing interviews and did not even have the chance to fully explore my options.</p>

<p>All in all, I am glad that things worked out so I did not have to deal with the funding issue (as I pushed back my graduation) and the anxiety of continuous job search (at least for the near term, hopefully). No words can describe how grateful I am for all the people who supported me during the process.</p>

<h3 id="surprise-1-only-one-or-two-papers-really-matter-during-my-job-search">Surprise 1: Only one or two papers really matter during my job search.</h3>

<p>Based on <a href="https://alisawuffles.github.io/blog/job-search/">Alisa’s post</a> and reactions, perhaps many already know that <strong>the interviews (e.g. LeetCode) may not be related to the research work you’ve done.</strong></p>

<p>I would take this even further and said that only one or two papers really matter during the job search. Sometimes, none at all, and I was just being evaluated on how well I solve the team’s problems on the spot.</p>

<p>In my experience: the roles of the papers are mostly two:</p>
<ol>
  <li><strong>Get my foot in the door.</strong> I have worked on something that the team liked, or my paper has demonstrated certain expertise the team is looking for, so I am now put into the interview pipeline. That is, I just passed the bar and now I’m officially being considered as an applicant.</li>
  <li><strong>Deep dive.</strong> This happens usually during research presentation or research discussion, where I talk about the motivation and the details behind one work. Sometimes, such presentation can be as short as 20 minutes only.</li>
</ol>

<p>So to some degree, the volume of publication does not really matter aside from establishing credibility. In my case, my multilingual research papers substantially outnumber my AI safety papers–––but given my pivot to AI safety research, none of those work have any bearing on my interview outcome, including papers I received best paper awards on.</p>

<p>This is actually <strong>liberating</strong>, because that means that you can always pivot to a new field that you think is impactful and still get dream offers if you demonstrate sufficient expertise in that field and the team wants you. On the flip side, it also suggests that you’d need to keep up with the field, as past success has less bearing on whether you’d get hired into new opportunities.</p>

<h3 id="surprise-2-very-diverse-interview-rounds">Surprise 2: Very diverse interview rounds.</h3>

<p>I originally came into the interviews expecting something like how fresh-grad software engineers being interviewed (e.g., Leetcode-style questions and behavioral rounds) plus some technical rounds about LLM/deep learning.</p>

<p>That there’s something standardized about the interview rounds–––for which I believe the blogs from Alisa and Silvia give the impression of.</p>

<p>Surprisingly, I have received questions about system design as well as parallel programming (such as using <code class="language-plaintext highlighter-rouge">asyncio</code> to implement concurrency operations) during my job search. I also learned that there are interview rounds where you are evaluated on how well you use AI agents. All in all, the lesson here is that you should always expect wildcard questions and diverse interview rounds.</p>

<h3 id="surprise-3-work-trials">Surprise 3: Work trials.</h3>

<p>This is something entirely new to me. It was also surprising for me when I saw that in Alisa’s post since I thought work trials are only common for AI safety positions. Apparently, it is increasingly common for AI startups as well.</p>

<p>Work trials are completely different from onsite–––you are not flown to the company to do multiple interview rounds onsite; instead, you are working with the team to solve a task. Sometimes, the task can be open-ended.</p>

<p>These work trials are paid usually, but what surprises me is that some of these in-person work trials can last up to a week.</p>

<p>For me, doing work trials make it really hard to prepare for other companies’ interviews as I would have to put in my everything on the current task assignment and have no bandwidth for interview prep with other companies. This is something you should be mindful of when you schedule interviews, especially if you are interviewing with multiple companies simultaneously and have tight turnaround time.</p>

<h3 id="surprise-4-timing-matters-a-lot">Surprise 4: Timing matters a lot.</h3>

<p>In this current job market, timing plays a substantial role.</p>

<p>For instance, in the last Fall, it was extremely challenging to find AI safety positions compared to positions related to RL. But now, there are more startups offering opportunities related to AI safety (such as Lila and Mechanize).</p>

<p>There are a few discussion points about how timing affects your search for full-time positions:</p>
<ol>
  <li>Your work got viral, and a lot of orgs take interest in your work and want to recruit you. You might get caught off guard by the timing, and the best thing you could do here is to take advantage of that timing and go through the interviews.</li>
  <li>Your area of research is becoming more popular. This is related to the AI safety example I mentioned above. You can assume that the opportunities are generally more available. The job application windows can be as short as under a month, or can span several months as the companies are trying to grow.</li>
  <li>Headcounts. This is something you should ask your recruiters about especially if you are planning to postpone interviews or doing some meta-planning about how to simultaneously interview with multiple companies.</li>
  <li>Exploding offers. If you were in this scenario, ask other companies to accelerate interviews. Don’t be surprised if you have to have three back-to-back interviews within a single day, and you only have less than a day to prepare for them.</li>
</ol>

<p>It is reasonable to ask to start the interviews later (like you can push till one or two months later), but usually once you have begun the interview, the intervals between each round are usually short. Another related note is that some positions expect you to start the role in the next month or two, though the start date can be negotiated.</p>

<h3 id="surprise-5-return-offers-are-rare">Surprise 5: Return offers are rare.</h3>

<p>Compared to software engineering positions, where return offers are usually a norm, for research roles it is more case-by-case.</p>

<p>For instance, during my Meta internship in 2024, return full-time offers were rare and highly headcount/team-dependent. Many of my friends did not get it. For my Astra fellowship with OpenAI, I still had to go through all the interview rounds to get into OpenAI just as any other applicant.</p>

<p>I heard that there are some other organizations where the interviews are more expedited; for instance, if team match works out, you just need to go through one or two more rounds.</p>

<h3 id="surprise-6-a-lot-of-interviews-are-not-your-topic-related">Surprise 6: A lot of interviews are not [your-topic]-related.</h3>

<p>This came as a surprise to me because I was pivoting from doing capability research (multilinguality) to safety, and I thought safety-related interviews would take a big portion of full interview pipeline. This impression was amplified by how much AI safety discussion was constantly held within Constellation during my Astra Fellowship.</p>

<p>That wasn’t true.</p>

<p>In fact, I have encountered many rounds not related to AI safety at all, let alone related to my research interest. I believe this experience is similar to what’s shared by Alisa and Silvia (even though they work on other AI fields).</p>

<p>In a handful of places, it still felt like I was evaluated on how well-rounded an AI researcher I was. I believe there’s merits to this (e.g., field is fast-moving so checking on fundamentals is important, etc.), but I was definitely expecting a higher ratio of AI-safety-related questions since it is in my opinion an urgent research topic, and it is still a rather niched field. Perhaps my interviewing experience might be more different for senior positions.</p>

<p><strong>For safety researchers:</strong> If it is helpful for you, I have co-written <a href="https://www.lesswrong.com/posts/dvsFfGuXXyHYkyifp/tips-for-cracking-the-ai-safety-technical-interview-1">a LessWrong post</a> about the safety-specific rounds, but expect a lot of diversity in the questions asked.</p>

<h3 id="reading-resources-i-recommend-2026">Reading Resources I recommend (2026):</h3>

<p>The following are resources I can vouch for on how to prepare for the job market.</p>

<ol>
  <li>Nathan Lambert – <a href="https://www.interconnects.ai/p/thoughts-on-the-hiring-market-in">Thoughts on the job market in the age of LLMs</a></li>
  <li>Alisa Liu – <a href="https://alisawuffles.github.io/blog/job-search/">Notes on the Industry Job Search</a></li>
  <li>Silvia Sapora – <a href="https://silviasapora.github.io/blog/ml-interviews.html">ML Job Interviews: The Ultimate Guide</a></li>
</ol>

<hr />

<div class="footnotes" role="doc-endnotes">
  <ol>
    <li id="fn:timing">
      <p>One of <a href="https://arxiv.org/abs/2604.03121">my (side) project</a> was well-received and received cold emails/DMs from hiring managers in April. <a href="#fnref:timing" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
  </ol>
</div>]]></content>
          
          <author>
            <name>Yong Zheng-Xin</name>
          </author>
          
            <category term="non-technical" />
          
          
          
            <summary type="html"><![CDATA[Things I wish I knew before.]]></summary>
          
        </entry>
        
      
    
  
    
      
      
      
        
        
        <entry>
          <title type="html">CoT monitorability: why you should use g-means over F1</title>
          <link href="https://yongzx.github.io/blog/2026/02/09/cot-monitorability-why-g-means-and-not-f1/?utm_source=rss&amp;utm_medium=feed&amp;utm_campaign=blog" rel="alternate" type="text/html" title="CoT monitorability: why you should use g-means over F1" />
          <published>2026-02-09T00:00:00-08:00</published>
          <updated>2026-02-09T00:00:00-08:00</updated>
          <id>https://yongzx.github.io/blog/2026/02/09/cot-monitorability-why-g-means-and-not-f1/</id>
          
            <content type="html" xml:base="https://yongzx.github.io/blog/2026/02/09/cot-monitorability-why-g-means-and-not-f1/"><![CDATA[<p>The Monitoring Monitorability paper <sup class="citation-marker"><a href="#ref-1" id="ref-1-cite-1">[1]</a></sup> argues that g-mean is a better metric than F1. In this blog post, I want to explain my understanding about it.</p>

<h2 id="preliminary-math">Preliminary math</h2>

<p>Suppose you have a binary monitor that says either “misbehavior detected” or “no misbehavior.” The ground truth is either positive (bad behavior happened) or negative (it did not). As usual, the four outcomes are TP, FP, TN, and FN. Let the total number of samples be $\mathcal{N}$.</p>

<p>We then have:</p>

<ul>
  <li>Prevalence $\pi$ asks how often does the model exhibit bad behaviors.</li>
</ul>

\[\pi = \frac{\mathrm{TP} + \mathrm{FN}}{\mathcal{N}}\]

<ul>
  <li>$\mathrm{TPR}$ means that how many bad behaviors did we flag out of all bad behaviors. This is also <em>recall</em>.</li>
</ul>

\[\mathrm{TPR} = \frac{\mathrm{TP}}{\mathrm{TP} + \mathrm{FN}}\]

<ul>
  <li>$\mathrm{TNR}$ means that how many good behaviors did we flag out of all good behaviors.</li>
</ul>

\[\mathrm{TNR} = \frac{\mathrm{TN}}{\mathrm{TN} + \mathrm{FP}}\]

<ul>
  <li>$\mathrm{Precision}$ flips the $\mathrm{TPR}$, which is now asking: out of all the bad behaviors we flag, how many are truly bad?</li>
</ul>

\[\mathrm{Precision} = \frac{\mathrm{TP}}{\mathrm{TP} + \mathrm{FP}}\]

<!--- Important part -->
<h2 id="key-difference-between-g-mean-and-f1">Key difference between g-mean and F1</h2>

<p>The key difference between g-mean and F1 is whether prevalence $\pi$ interferes with the metric. First, let us break down g-mean:</p>

\[\mathrm{g\text{-}mean} = \sqrt{\mathrm{TPR} \cdot \mathrm{TNR}} = \sqrt{\frac{\text{flagged-bad}}{\text{all-bad}} \cdot \frac{\text{flagged-good}}{\text{all-good}}}.\]

<p>TPR draws both its numerator and denominator from the (all-bad) positive pool, i.e. the samples where misbehavior actually happened. TNR draws both from the (all-good) negative pool, i.e. the samples where no misbehavior happened. Because TPR and TNR only look <em>within</em> their own class, changing the relative size of the two pools––meaning if you change $\pi$––does not affect TPR and TNR.</p>

<p>Now, let us express F1 with TPR and TNR:</p>

\[\mathrm{F1} = \frac{2 \cdot \mathrm{Precision} \cdot \mathrm{Recall}}{\mathrm{Precision} + \mathrm{Recall}} = \frac{2 \cdot \mathrm{Precision} \cdot \mathrm{TPR}}{\mathrm{Precision} + \mathrm{TPR}} .\]

<p>The blocker is Precision, which we need to bring in $\pi$.</p>

\[\mathrm{TP} = \mathrm{TPR} \cdot \pi \mathcal{N}\]

\[\mathrm{FP} = (1 - \mathrm{TNR}) \cdot (1 - \pi)\mathcal{N}.\]

\[\mathrm{Precision}
= \frac{\mathrm{TP}}{\mathrm{TP} + \mathrm{FP}}
= \frac{\mathrm{TPR}\pi}{\mathrm{TPR}\pi + (1-\mathrm{TNR})(1-\pi)}.\]

\[\mathrm{F1}
= \frac{2\pi \cdot \mathrm{TPR}}
{\pi(1+\mathrm{TPR}) + (1-\pi)(1-\mathrm{TNR})}.\]

<p>This is the crucial point: F1 explicitly depends on $\pi$, while g-mean does not.</p>

<!-- new subsection-->
<h2 id="so-what-is-wrong-with-prevalence-pi">So what is wrong with prevalence $\pi$?</h2>

<p>The key idea behind g-mean is that we do <strong>not</strong> want the monitorability metric to be confounded by prevalence $\pi$, i.e. by how likely the model is to misbehave.</p>

<p>Each agent has two independent properties in a monitorability setup:</p>

<ul>
  <li>Behavioral tendency ($\pi$): how often it misbehaves on these tasks.</li>
  <li>CoT transparency: when it misbehaves (bad behavior), how often does the bad behavior get flagged; and when it does not (good behavior), similarly how often does it get identified.</li>
</ul>

<p>The problem with F1 is that it includes $\pi$ through precision, so it collapses these two different properties into one number. For g-means, it is just the latter property.</p>

<h3 id="a-simple-thought-experiment">A simple thought experiment</h3>

<p>Imagine the paper runs an impossible coding task on five models, all monitored by the same GPT-5 monitor. Suppose every model’s chain of thought is equally transparent, and the monitor achieves $\mathrm{TPR} = 0.99$ and $\mathrm{TNR} = 0.99$ on all of them. In other words, they are all equally monitorable. A correct metric should therefore give the same score to all five.</p>

<p>But suppose the models misbehave at different rates:</p>

<ul>
  <li>Model A hacks 90% of the time, so $\pi = 0.9$.</li>
  <li>Model B hacks 50% of the time, so $\pi = 0.5$.</li>
  <li>Model C hacks 20% of the time, so $\pi = 0.2$.</li>
  <li>Model D hacks 5% of the time, so $\pi = 0.05$.</li>
  <li>Model E hacks 1% of the time, so $\pi = 0.01$.</li>
</ul>

<p>Now imagine publishing the F1 scores. You will get something like:</p>

<ul>
  <li>Model E monitorability: 0.33 (penalized by low prevalence $\pi$)</li>
  <li>Model A monitorability: 0.99</li>
</ul>

<p><img src="/assets/posts/cot-monitorability/case-b-plot.png" alt="Case B: TPR = 99%, TNR = 99%" /></p>
<p class="post-caption">Figure 3 (Case B) taken from <span class="citation-inline"><a href="#ref-1" id="ref-1-cite-2">[1]</a></span>, where Model E with low prevalence (high fraction of negatives) corresponds to the major drop of blue-line.</p>

<p>The natural conclusion would be: “Model E’s CoT is nearly opaque; we can barely monitor it; this is dangerous.”</p>

<p>But that conclusion is false. Model E barely misbehaves, and when it does, the monitor catches it just as reliably as it catches Model A. F1 makes the safest model look the most dangerous because it entangles two independent facts:</p>

<ul>
  <li>how often the model misbehaves</li>
  <li>how observable that misbehavior is</li>
</ul>

<p>In reality, Model E is the safest model in the lineup for two independent reasons: it rarely misbehaves <strong>and</strong> it is transparent when it does.</p>

<!-- Dsicussion -->
<h2 id="discussion">Discussion</h2>

<h3 id="what-is-special-about-g-mean">What is special about g-mean?</h3>
<p>Nothing uniquely magical. Any metric that symmetrically aggregates TPR and TNR without introducing prevalence would do. G-mean is simply the cleanest choice and strongly penalizes failure on either axis.</p>

<h3 id="limitations-of-g-mean">Limitations of g-mean</h3>
<p>That said, g-mean does have limitations. Since it excludes $\pi$, it is also insensitive to the balance between TPR and TNR. For example, a monitor with</p>

\[\mathrm{TPR} = 0.64, \qquad \mathrm{TNR} = 1.0\]

<p>has the same g-mean as a monitor with</p>

\[\mathrm{TPR} = \mathrm{TNR} = 0.8.\]

<p>In a safety-critical setting, those monitors are not equivalent. The first misses 36% of actual misbehaviors, which could be unacceptable even if it never raises a false alarm.</p>

<h2 id="appendix-degenerate-monitor-that-always-flag-bad">Appendix: Degenerate monitor that always flag “bad”</h2>

<p>Consider a degenerate monitor that always predicts positive. Then</p>

<ul>
  <li>$\mathrm{TPR} = 1$ because it catches all positives.</li>
  <li>$\mathrm{TNR} = 0$ because it never correctly identifies negatives.</li>
</ul>

<p>Using the earlier expression for F1,</p>

\[\mathrm{F1}
= \frac{2\pi \cdot \mathrm{TPR}}{\pi(1+\mathrm{TPR}) + (1-\pi)(1-\mathrm{TNR})}.\]

<p>Substituting $\mathrm{TPR}=1$ and $\mathrm{TNR}=0$ gives</p>

\[\mathrm{F1}
= \frac{2\pi}{\pi(1+1) + (1-\pi)(1-0)}
= \frac{2\pi}{\pi + 1}.\]

<p>If prevalence is extremely high, say $\pi = 0.999$, then</p>

\[\mathrm{F1} = \frac{2(0.999)}{1 + 0.999} \approx 0.999.\]

<p>So this degenerate monitor would be classified as nearly perfect, even though it is obviously useless on negatives.</p>

<p>By contrast, g-mean gives</p>

\[\sqrt{\mathrm{TPR} \cdot \mathrm{TNR}} = \sqrt{1 \cdot 0} = 0.\]

<p>This issue is partly about the numerator in F1 using only TPR, but prevalence is the main reason the score becomes misleadingly high for degenerate monitors. A simple alternative would be to use something proportional to $\mathrm{TPR} \cdot \mathrm{TNR}$ instead.</p>

<h2 id="acknowledgements">Acknowledgements</h2>

<p>Before release, I received helpful feedback from Miles Wang on the post’s explanation of g-mean.</p>

<h2 id="references">References</h2>
<ol class="bibliography">
<li id="ref-1">
  <span class="reference-authors">Guan, Melody Y., Wang, Miles, Carroll, Micah, et al.</span>
  <em>Monitoring Monitorability</em>.
  <span class="reference-venue">arXiv preprint, 2025.</span>
  <a href="https://arxiv.org/abs/2512.18311">arXiv:2512.18311</a>
  <a class="reference-backlink" href="#ref-1-cite-1" aria-label="Jump back to citation 1">↩1</a> <a class="reference-backlink" href="#ref-1-cite-2" aria-label="Jump back to citation 2">↩2</a>
</li>
</ol>]]></content>
          
          <author>
            <name>Yong Zheng-Xin</name>
          </author>
          
          
          
            <summary type="html"><![CDATA[It all has to do with prevalence.]]></summary>
          
        </entry>
        
      
    
  
    
      
      
      
        
        
        <entry>
          <title type="html">On burnout as a PhD student</title>
          <link href="https://yongzx.substack.com/p/on-burnout" rel="alternate" type="text/html" title="On burnout as a PhD student" />
          <published>2026-02-03T00:00:00-08:00</published>
          <updated>2026-02-03T00:00:00-08:00</updated>
          <id>https://yongzx.substack.com/p/on-burnout</id>
          
          <author>
            <name>Yong Zheng-Xin</name>
          </author>
          
            <category term="non-technical" />
          
          
          
            <summary type="html"><![CDATA[A short reflection on what actually helped avoid personal burnout during the PhD.]]></summary>
          
        </entry>
        
      
    
  
    
      
      
      
        
        
        <entry>
          <title type="html">RL vs next-token prediction: why a dichotomy?</title>
          <link href="https://yongzx.substack.com/p/rl-vs-next-token-prediction-why-should" rel="alternate" type="text/html" title="RL vs next-token prediction: why a dichotomy?" />
          <published>2025-07-11T00:00:00-07:00</published>
          <updated>2025-07-11T00:00:00-07:00</updated>
          <id>https://yongzx.substack.com/p/rl-vs-next-token-prediction-why-should</id>
          
          <author>
            <name>Yong Zheng-Xin</name>
          </author>
          
            <category term="technical" />
          
          
          
            <summary type="html"><![CDATA[Conceptual note on why RL and NTP training should be merged.]]></summary>
          
        </entry>
        
      
    
  
    
      
      
      
        
        
        <entry>
          <title type="html">Shuchao Bi – Advancing the Frontier of Silicon Intelligence: the Past, Open Problems, and the Future</title>
          <link href="https://yongzx.substack.com/p/shuchao-bi-advancing-the-frontier" rel="alternate" type="text/html" title="Shuchao Bi – Advancing the Frontier of Silicon Intelligence: the Past, Open Problems, and the Future" />
          <published>2025-07-01T00:00:00-07:00</published>
          <updated>2025-07-01T00:00:00-07:00</updated>
          <id>https://yongzx.substack.com/p/shuchao-bi-advancing-the-frontier</id>
          
          <author>
            <name>Yong Zheng-Xin</name>
          </author>
          
            <category term="notes" />
          
          
          
        </entry>
        
      
    
  
    
      
      
      
        
        
        <entry>
          <title type="html">Emergent misalignment vs jailbreaks: a short analysis</title>
          <link href="https://yongzx.substack.com/p/emergent-misalignment-vs-jailbroken" rel="alternate" type="text/html" title="Emergent misalignment vs jailbreaks: a short analysis" />
          <published>2025-06-20T00:00:00-07:00</published>
          <updated>2025-06-20T00:00:00-07:00</updated>
          <id>https://yongzx.substack.com/p/emergent-misalignment-vs-jailbroken</id>
          
          <author>
            <name>Yong Zheng-Xin</name>
          </author>
          
            <category term="technical" />
          
          
          
            <summary type="html"><![CDATA[A short note on the differences between emergent misalignment and jailbroken models.]]></summary>
          
        </entry>
        
      
    
  
    
      
      
      
        
        
        <entry>
          <title type="html">AI research internship hunt (2023) as a CS PhD student</title>
          <link href="https://yongzx.substack.com/p/ai-research-internship-search-as" rel="alternate" type="text/html" title="AI research internship hunt (2023) as a CS PhD student" />
          <published>2024-03-03T00:00:00-08:00</published>
          <updated>2024-03-03T00:00:00-08:00</updated>
          <id>https://yongzx.substack.com/p/ai-research-internship-search-as</id>
          
          <author>
            <name>Yong Zheng-Xin</name>
          </author>
          
            <category term="non-technical" />
          
          
          
            <summary type="html"><![CDATA[Notes from a relatively successful search for summer research internships as a third-year PhD student.]]></summary>
          
        </entry>
        
      
    
  
</feed>
