<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en"><generator uri="https://jekyllrb.com/" version="4.4.1">Jekyll</generator><link href="https://asjad99.github.io/feed.xml" rel="self" type="application/atom+xml"/><link href="https://asjad99.github.io/" rel="alternate" type="text/html" hreflang="en"/><updated>2026-08-23T03:53:48+00:00</updated><id>https://asjad99.github.io/feed.xml</id><title type="html">Asjad K. (PhD)</title><subtitle>Data scientist, researcher, and builder exploring how intelligent systems create value in the real world. </subtitle><entry><title type="html">Navigating Complexity: The GE-McKinsey Nine-Box Matrix</title><link href="https://asjad99.github.io/blog/2026/04/navigating-complexity-the-ge-mckinsey-nine-box-matrix-2/" rel="alternate" type="text/html" title="Navigating Complexity: The GE-McKinsey Nine-Box Matrix"/><published>2026-04-13T00:01:36+00:00</published><updated>2026-04-13T00:01:36+00:00</updated><id>https://asjad99.github.io/blog/2026/04/navigating-complexity-the-ge-mckinsey-nine-box-matrix-2</id><content type="html" xml:base="https://asjad99.github.io/blog/2026/04/navigating-complexity-the-ge-mckinsey-nine-box-matrix-2/"><![CDATA[<p>The proliferation of multibusiness corporations in the 20th century posed unprecedented challenges to effective strategic management, particularly with regard to resource allocation across diverse business units. As companies expanded and diversified, they required robust frameworks to evaluate and prioritize investments among disparate units. The GE-McKinsey nine-box matrix, introduced in the early 1970s, emerged as a sophisticated solution to these challenges, enabling corporations like General Electric to make informed decisions on where to invest, which units to grow, and which to divest. This framework became a critical tool for addressing the complexity of managing a diversified portfolio of business units, providing a systematic approach to ensure that corporate resources were allocated efficiently and strategically to maximize overall corporate value.</p> <p><strong>The Origin and Purpose of the Nine-Box Matrix</strong></p> <p>The GE-McKinsey nine-box matrix was conceived as an evolution of the Boston Consulting Group's (BCG) growth-share matrix—a framework that had garnered widespread use for assessing the balance between market growth and market share within business units. The nine-box matrix sought to refine and extend the BCG framework by addressing the needs of large, decentralized organizations with greater precision. Unlike its predecessor, which focused solely on market growth and relative market share, the GE-McKinsey matrix introduced a more nuanced approach by incorporating additional dimensions of strategic evaluation. This allowed for a deeper understanding of each business unit's potential, enabling more tailored strategic actions.</p> <p>At its core, the GE-McKinsey nine-box matrix evaluates business units based on two critical dimensions: <strong>industry attractiveness</strong> and <strong>competitive strength</strong>. These dimensions collectively determine the unit's placement within the nine-cell grid, which offers an analytical foundation for strategic resource allocation. The matrix goes beyond simplistic categorizations, offering a structured method to understand the complexities of different business units and their respective potential for contributing to the corporation's overall success.</p> <p><strong>Industry Attractiveness and Competitive Strength</strong></p> <p>The nine-box matrix requires a rigorous assessment of <strong>industry attractiveness</strong>, which encapsulates factors such as market size, growth rate, profitability, barriers to entry, and the potential for technological or product innovation. Industry attractiveness is a multifaceted construct that considers both the present conditions and future potential of an industry, allowing corporations to assess whether a particular sector is worth continued investment or represents a declining opportunity. Evaluating <strong>competitive strength</strong> involves a parallel analysis of a business unit's positioning within its industry, taking into account metrics such as market share, brand equity, operational efficiency, distribution capabilities, and product differentiation. Competitive strength assessment involves understanding both quantitative and qualitative factors, which helps determine the degree of control and influence a business unit can exert within its market.</p> <p>The intersection of these two dimensions provides a framework for identifying where a business unit resides within the matrix. Units that combine high industry attractiveness with high competitive strength represent prime candidates for significant investment. These units are often the core drivers of growth and profitability for the corporation, and therefore warrant focused efforts to expand their capabilities and market reach. Conversely, units characterized by low competitive strength in unattractive industries may warrant divestiture or strategic repositioning aimed at cash generation. Such decisions require careful analysis, as divesting from a unit could also impact related parts of the business, particularly if synergies exist between units.</p> <p><strong>Using the Matrix to Drive Strategy</strong></p> <p>The placement of business units within the nine-box matrix provides an <strong>analytic portfolio management tool</strong> for executives to navigate the complexities of a multibusiness corporation. Business units are classified into three strategic categories, each of which has distinct implications for how resources should be allocated:</p> <ol><li><strong>Grow and Invest</strong>: Business units positioned in the upper right quadrant of the matrix typically operate in attractive industries while possessing a strong competitive advantage. These units are prime targets for resource allocation aimed at driving growth and expanding market leadership. Investments in these units are often focused on scaling operations, enhancing product offerings, and expanding market presence. The objective is to capitalize on existing strengths while solidifying the unit's leadership position in a high-growth market.</li><li><strong>Selectively Invest</strong>: Units along the matrix's diagonal represent moderate industry attractiveness and competitive strength. These units may warrant selective investment, with the aim of gradually enhancing their competitiveness while remaining mindful of emerging opportunities or risks. These business units often require a balanced approach—investments are made to maintain or slightly improve market position, while also being cautious about not overcommitting resources. The objective is to identify pathways to growth while managing risk and avoiding significant exposure to volatile or uncertain market conditions.</li><li><strong>Harvest or Divest</strong>: Units occupying the lower left quadrant, characterized by low attractiveness and limited competitive strength, are better suited for harvesting—maximizing short-term cash flows with minimal reinvestment—or divestiture to free up resources for more promising ventures. Harvesting involves a focus on cost-cutting and efficiency to extract as much value as possible, while divestiture involves finding suitable buyers who may see potential in the unit. These strategic choices must be made with a clear understanding of the broader corporate strategy, ensuring that divesting underperforming units does not inadvertently weaken the overall portfolio.</li></ol> <p><strong>Balancing the Portfolio</strong></p> <p>The true value of the nine-box matrix lies not only in categorizing business units but also in <strong>optimizing the overall corporate portfolio</strong>. A well-balanced portfolio typically comprises a mix of high-growth units, stable cash-generating units, and high-risk ventures with significant upside potential. The matrix facilitates strategic decision-making by providing a systematic framework for aligning corporate investments with long-term strategic goals. This process involves not only identifying which units require more resources but also determining how the mix of business units can achieve the desired balance between risk and return.</p> <p>Effective portfolio management also involves assessing the interdependencies among business units. For example, a high-growth unit may benefit from the cash flow generated by a mature, stable unit. By viewing the corporation as an interconnected system of businesses, executives can make more informed decisions that leverage synergies and optimize resource allocation. The nine-box matrix serves as a powerful visual representation of these dynamics, allowing executives to identify gaps, redundancies, and opportunities for cross-business collaboration.</p> <p>Importantly, the utility of the nine-box matrix extends beyond rigid categorization—it requires <strong>executive judgment</strong> and a deep understanding of contextual dynamics. For instance, a business unit that demonstrates strong competitive positioning within a declining industry presents a complex scenario that necessitates careful deliberation on whether to invest in innovation or divest and reallocate resources to more promising sectors. Strategic decisions in such scenarios require a nuanced approach, where quantitative analysis is complemented by qualitative insights into market trends, customer behavior, and potential disruptions.</p> <p><strong>The Legacy and Evolution of the Nine-Box Framework</strong></p> <p>The GE-McKinsey nine-box matrix laid the groundwork for numerous subsequent portfolio models, such as the Matrix for Assessing Corporate Strategies (MACS) and the Portfolio of Initiatives. Over time, the criteria for evaluating industry attractiveness and competitive strength have become increasingly sophisticated, encompassing considerations such as environmental sustainability, digital disruption, and geopolitical risk. Modern adaptations of the matrix also incorporate advanced data analytics to assess factors such as customer sentiment, competitive threats, and macroeconomic trends in real time.</p> <p>Despite its origins in the early 1970s, the nine-box matrix—and its derivatives—remains highly relevant in contemporary strategic management. Many large enterprises continue to rely on portfolio models inspired by this framework to determine optimal resource allocation and prioritize investments across an increasingly complex array of opportunities. The nine-box matrix has also influenced the development of other strategic tools that emphasize a holistic, data-driven approach to managing diversified portfolios.</p> <p>The continued relevance of the nine-box matrix can also be attributed to its adaptability. As industries evolve and new challenges arise, the framework has been modified to account for the changing nature of competition, technological advances, and shifts in consumer preferences. The matrix's foundational principles of assessing industry attractiveness and competitive strength remain constant, but the metrics and considerations used to evaluate these dimensions have evolved significantly. This adaptability has ensured that the nine-box matrix remains a cornerstone of strategic management in multibusiness corporations.</p> <p><strong>Conclusion</strong></p> <p>The GE-McKinsey nine-box matrix endures as a seminal framework for addressing the complexities inherent in multibusiness corporations. It provides a structured, yet flexible, methodology for making critical strategic decisions regarding investment and divestiture. By integrating systematic analysis with executive judgment, the matrix enables organizations to prioritize resources effectively, fostering the growth of their most promising business units while mitigating exposure to less viable ones. The matrix serves as both a diagnostic tool and a strategic compass, guiding corporate leaders in their pursuit of sustainable, long-term value creation.</p> <p>Ultimately, the continued relevance of the nine-box matrix highlights a fundamental lesson for today's leaders: effective strategic decision-making requires both analytical rigor and the ability to adapt to evolving circumstances. The interplay between systematic frameworks and nuanced managerial insights has always been, and will continue to be, the hallmark of successful strategic leadership. As corporations face an increasingly dynamic and complex business environment, the principles embodied by the nine-box matrix remain vital for navigating uncertainty, optimizing resource allocation, and achieving sustained competitive advantage.</p>]]></content><author><name></name></author><summary type="html"><![CDATA[The proliferation of multibusiness corporations in the 20th century posed unprecedented challenges to effective strategic management, particularly with regard to resource allocation across diverse business units. As companies expanded and diversified, they required robust frameworks to evaluate and prioritize investments among disparate units. The GE-McKinsey nine-box matrix, introduced in the early 1970s, emerged as a sophisticated solution to these challenges, enabling corporations like General Electric to make informed decisions on where to invest, which units to grow, and which to divest. This framework became a critical tool for addressing the complexity of managing a diversified portfolio of business units, providing a systematic approach to ensure that corporate resources were allocated efficiently and strategically to maximize overall corporate value. The Origin and Purpose of the Nine-Box Matrix The GE-McKinsey nine-box matrix was conceived as an evolution of the Boston Consulting Group's (BCG) growth-share matrix—a framework that had garnered widespread use for assessing the balance between market growth and market share within business units. The nine-box matrix sought to refine and extend the BCG framework by addressing the needs of large, decentralized organizations with greater precision. Unlike its predecessor, which focused solely on market growth and relative market share, the GE-McKinsey matrix introduced a more nuanced approach by incorporating additional dimensions of strategic evaluation. This allowed for a deeper understanding of each business unit's potential, enabling more tailored strategic actions. At its core, the GE-McKinsey nine-box matrix evaluates business units based on two critical dimensions: industry attractiveness and competitive strength. These dimensions collectively determine the unit's placement within the nine-cell grid, which offers an analytical foundation for strategic resource allocation. The matrix goes beyond simplistic categorizations, offering a structured method to understand the complexities of different business units and their respective potential for contributing to the corporation's overall success. Industry Attractiveness and Competitive Strength The nine-box matrix requires a rigorous assessment of industry attractiveness, which encapsulates factors such as market size, growth rate, profitability, barriers to entry, and the potential for technological or product innovation. Industry attractiveness is a multifaceted construct that considers both the present conditions and future potential of an industry, allowing corporations to assess whether a particular sector is worth continued investment or represents a declining opportunity. Evaluating competitive strength involves a parallel analysis of a business unit's positioning within its industry, taking into account metrics such as market share, brand equity, operational efficiency, distribution capabilities, and product differentiation. Competitive strength assessment involves understanding both quantitative and qualitative factors, which helps determine the degree of control and influence a business unit can exert within its market. The intersection of these two dimensions provides a framework for identifying where a business unit resides within the matrix. Units that combine high industry attractiveness with high competitive strength represent prime candidates for significant investment. These units are often the core drivers of growth and profitability for the corporation, and therefore warrant focused efforts to expand their capabilities and market reach. Conversely, units characterized by low competitive strength in unattractive industries may warrant divestiture or strategic repositioning aimed at cash generation. Such decisions require careful analysis, as divesting from a unit could also impact related parts of the business, particularly if synergies exist between units. Using the Matrix to Drive Strategy The placement of business units within the nine-box matrix provides an analytic portfolio management tool for executives to navigate the complexities of a multibusiness corporation. Business units are classified into three strategic categories, each of which has distinct implications for how resources should be allocated: Grow and Invest: Business units positioned in the upper right quadrant of the matrix typically operate in attractive industries while possessing a strong competitive advantage. These units are prime targets for resource allocation aimed at driving growth and expanding market leadership. Investments in these units are often focused on scaling operations, enhancing product offerings, and expanding market presence. The objective is to capitalize on existing strengths while solidifying the unit's leadership position in a high-growth market.Selectively Invest: Units along the matrix's diagonal represent moderate industry attractiveness and competitive strength. These units may warrant selective investment, with the aim of gradually enhancing their competitiveness while remaining mindful of emerging opportunities or risks. These business units often require a balanced approach—investments are made to maintain or slightly improve market position, while also being cautious about not overcommitting resources. The objective is to identify pathways to growth while managing risk and avoiding significant exposure to volatile or uncertain market conditions.Harvest or Divest: Units occupying the lower left quadrant, characterized by low attractiveness and limited competitive strength, are better suited for harvesting—maximizing short-term cash flows with minimal reinvestment—or divestiture to free up resources for more promising ventures. Harvesting involves a focus on cost-cutting and efficiency to extract as much value as possible, while divestiture involves finding suitable buyers who may see potential in the unit. These strategic choices must be made with a clear understanding of the broader corporate strategy, ensuring that divesting underperforming units does not inadvertently weaken the overall portfolio. Balancing the Portfolio The true value of the nine-box matrix lies not only in categorizing business units but also in optimizing the overall corporate portfolio. A well-balanced portfolio typically comprises a mix of high-growth units, stable cash-generating units, and high-risk ventures with significant upside potential. The matrix facilitates strategic decision-making by providing a systematic framework for aligning corporate investments with long-term strategic goals. This process involves not only identifying which units require more resources but also determining how the mix of business units can achieve the desired balance between risk and return. Effective portfolio management also involves assessing the interdependencies among business units. For example, a high-growth unit may benefit from the cash flow generated by a mature, stable unit. By viewing the corporation as an interconnected system of businesses, executives can make more informed decisions that leverage synergies and optimize resource allocation. The nine-box matrix serves as a powerful visual representation of these dynamics, allowing executives to identify gaps, redundancies, and opportunities for cross-business collaboration. Importantly, the utility of the nine-box matrix extends beyond rigid categorization—it requires executive judgment and a deep understanding of contextual dynamics. For instance, a business unit that demonstrates strong competitive positioning within a declining industry presents a complex scenario that necessitates careful deliberation on whether to invest in innovation or divest and reallocate resources to more promising sectors. Strategic decisions in such scenarios require a nuanced approach, where quantitative analysis is complemented by qualitative insights into market trends, customer behavior, and potential disruptions. The Legacy and Evolution of the Nine-Box Framework The GE-McKinsey nine-box matrix laid the groundwork for numerous subsequent portfolio models, such as the Matrix for Assessing Corporate Strategies (MACS) and the Portfolio of Initiatives. Over time, the criteria for evaluating industry attractiveness and competitive strength have become increasingly sophisticated, encompassing considerations such as environmental sustainability, digital disruption, and geopolitical risk. Modern adaptations of the matrix also incorporate advanced data analytics to assess factors such as customer sentiment, competitive threats, and macroeconomic trends in real time. Despite its origins in the early 1970s, the nine-box matrix—and its derivatives—remains highly relevant in contemporary strategic management. Many large enterprises continue to rely on portfolio models inspired by this framework to determine optimal resource allocation and prioritize investments across an increasingly complex array of opportunities. The nine-box matrix has also influenced the development of other strategic tools that emphasize a holistic, data-driven approach to managing diversified portfolios. The continued relevance of the nine-box matrix can also be attributed to its adaptability. As industries evolve and new challenges arise, the framework has been modified to account for the changing nature of competition, technological advances, and shifts in consumer preferences. The matrix's foundational principles of assessing industry attractiveness and competitive strength remain constant, but the metrics and considerations used to evaluate these dimensions have evolved significantly. This adaptability has ensured that the nine-box matrix remains a cornerstone of strategic management in multibusiness corporations. Conclusion The GE-McKinsey nine-box matrix endures as a seminal framework for addressing the complexities inherent in multibusiness corporations. It provides a structured, yet flexible, methodology for making critical strategic decisions regarding investment and divestiture. By integrating systematic analysis with executive judgment, the matrix enables organizations to prioritize resources effectively, fostering the growth of their most promising business units while mitigating exposure to less viable ones. The matrix serves as both a diagnostic tool and a strategic compass, guiding corporate leaders in their pursuit of sustainable, long-term value creation. Ultimately, the continued relevance of the nine-box matrix highlights a fundamental lesson for today's leaders: effective strategic decision-making requires both analytical rigor and the ability to adapt to evolving circumstances. The interplay between systematic frameworks and nuanced managerial insights has always been, and will continue to be, the hallmark of successful strategic leadership. As corporations face an increasingly dynamic and complex business environment, the principles embodied by the nine-box matrix remain vital for navigating uncertainty, optimizing resource allocation, and achieving sustained competitive advantage.]]></summary></entry><entry><title type="html">US Economy and Why Nations fail</title><link href="https://asjad99.github.io/blog/2026/04/us-economy-and-why-nations-fail/" rel="alternate" type="text/html" title="US Economy and Why Nations fail"/><published>2026-04-12T23:57:40+00:00</published><updated>2026-04-12T23:57:40+00:00</updated><id>https://asjad99.github.io/blog/2026/04/us-economy-and-why-nations-fail</id><content type="html" xml:base="https://asjad99.github.io/blog/2026/04/us-economy-and-why-nations-fail/"><![CDATA[<h3 id="what-are-the-structural-problems-that-the-trump-administration-is-focused-on-solving-what-is-this-administration%E2%80%99s-vision-for-america-what-are-the-implications-for-entrepreneurs-and-investors">What are the structural problems that the Trump administration is focused on solving? What is this administration’s vision for America? What are the implications for entrepreneurs and investors?</h3> <p></p> <p>The international geopolitical world order is breaking down. Up until now, monetary/economic order dictated that countries like China manufacture inexpensively, sell to Americans, and acquire American debt assets. But we are entering an new era of de-globalization, where rules of monetary, political, and geopolitical orders are being rewritten at an unprecedented rate.</p> <p>This article attempts to understand the biggest disruptions that are likely still ahead. But first we awnser, What are the underlying circumstances that caused US citizens to first elect Donald Trump and then Trump to implement massive tariffs, introduce a DOGE program, adjust taxes and so on.,</p> <p>The United States' national debt currently over $34 trillion is the result of a combination of historical, structural, and policy-driven factors that have accumulated over decades. US Public Debt Per Capita is at a current level of 106.11K, up from 104.06K last month and up from 100.41K one year ago. This level of current debt level and the rate at which the government debt is being added to is unsustainable.</p> <p></p> <figure class="kg-card kg-image-card"><img src="/assets/img/ghost/128902b40d-https_3A_2F_2Fsubstack-post-media.s3.amazonaws.com_2Fpublic_2Fimages_2Fe6e01ed5-6428-44d8-a0b2-490be3f6040a_4442x2327.webp" class="kg-image" alt="National debt of the United States - Wikipedia" loading="lazy" title="National debt of the United States - Wikipedia" width="1456" height="763"/></figure> <p></p> <h3 id="how-did-the-us-get-here-here"><strong>How Did the US Get here here?</strong></h3> <p></p> <p>Two third of American population is not college educated. Historically upward mobility was possible because there we lots of opportunities to find trade and blue colar jobs. But in recent decades, cheaper Chinese manufacturing - led to deteriorating of American manufacturing, which both hollows out middle class jobs in the U.S. and requires America to import needed items from a country that it is increasingly seeing as an enemy. It also leads to stagnation of wages for the white color jobs. All of this ultimately leads to<strong> breakdown of domestic political order</strong> due to huge gaps in people's education levels, opportunity levels, productivity levels, income and wealth levels, and values. History also shows that strong autocratic leaders emerge as classic democracy and classic rule of law are removed as barriers to autocratic leadership.</p> <p>At the moment, the U.S. runs a <strong>budget deficit</strong> when its annual spending exceeds its revenue. These deficits are financed by borrowing (i.e., issuing Treasury securities). Over time, recurring deficits accumulate into large national debt. Even in good economic times, the U.S. often spends more than it collects due to entitlement programs and defense spending. Furthermore, Tax cuts, such as those in 2001 (Bush-era) and 2017 (Trump-era), reduced federal revenue without equivalent spending cuts.</p> <p>In summary, Chronic budget deficits driven by entitlement spending, tax cuts, and military spending have been a driving force with interest rate projected to compound each year until the whole system is no longer sustainable and collapses (like the many empires of the past).</p> <p>The biggest components of U.S. government spending are, Social Security, Medicare, Defense and<strong> Interest on debt !</strong></p> <figure class="kg-card kg-image-card"><img src="/assets/img/ghost/76e4da48ae-https_3A_2F_2Fsubstack-post-media.s3.amazonaws.com_2Fpublic_2Fimages_2Fa6a07d26-112b-4add-8a50-b61120e98db0_1336x622.png" class="kg-image" alt="" loading="lazy" width="1336" height="622"/></figure> <p>Many of these programs are politically difficult to cut and are growing due to <strong>demographic shifts</strong>, especially the aging Baby Boomer population (more retirees = more Medicare/Social Security spending).</p> <p>The U.S. has borrowed heavily in response to major events:</p> <ul><li><strong>2001–2020 wars</strong> (Afghanistan, Iraq): trillions spent, mostly borrowed.</li><li><strong>2008 Global Financial Crisis</strong>: stimulus packages and bailouts.</li><li><strong>2020 COVID-19 pandemic</strong>: unprecedented $5+ trillion in relief spending.</li></ul> <p>These crises were partly unavoidable, but they contributed massively to the debt.</p> <h3 id="interest-rates-and-politics"><strong>Interest Rates and Politics:</strong></h3> <p>For more than a decade, US has been in a unique global position that has allowed it to borrow more easily than other nations. Additionally, U.S. borrowed at <strong>historically low interest rates</strong>, which made debt accumulation seem sustainable. Policymakers became less cautious about adding to the debt, as the <strong>cost of servicing it</strong> was low.</p> <figure class="kg-card kg-image-card"><img src="/assets/img/ghost/869eddbb8d-https_3A_2F_2Fsubstack-post-media.s3.amazonaws.com_2Fpublic_2Fimages_2F1f3f04a6-ce9b-4036-907c-7efde2725fb1_1695x944.png" class="kg-image" alt="" loading="lazy" width="1456" height="811"/></figure> <p>There always has been a political reluctance to raise taxes or cut popular programs as the political operation system encourages short-term decision-making.</p> <p>This has now changed: <strong>interest payments</strong> are the <strong>fastest-growing part of the federal budget</strong> due to rising rates.</p> <p>The seasoned Politicians can’t seem to do much about it due to political Gridlock and Short-Term Incentives.</p> <ul><li><strong>No long-term fiscal plan</strong>: U.S. politicians often avoid unpopular decisions like raising taxes or cutting benefits.</li><li><strong>Debt ceiling fights</strong> don’t reduce debt—they only delay or complicate borrowing.</li><li><strong>Populist pressure</strong>: There’s little political reward for fiscal discipline, especially when voters demand both low taxes and generous public programs.</li></ul> <hr/> <p></p> <p></p> <h2 id="balancing-the-budget">Balancing the Budget:</h2> <p>There are two forces to worry about growth and inflation. and therefore, balancing a budget without brining along the congress:</p> <ul><li>Cut Government spending (DOGE)</li><li>Earn more</li></ul> <h3 id="understanding-the-tarrif-policy"><strong>Understanding the Tarrif Policy:</strong></h3> <p><strong>Here is one theory explaining trump Govt. behaviour:</strong></p> <blockquote>During a recent MSNBC segment, Rachel Maddow unpacked the bizarre origin of Donald Trump’s trade policy during his first term. In 2016, seeking to bolster Trump’s economic team, Jared Kushner searched Amazon for books on China. He came across Death by China, co-authored by a little-known UC Irvine professor named Peter Navarro. Impressed by its aggressive tone, Kushner contacted Navarro, who was soon advising Trump directly on trade policy.<br/><br/>Navarro became the architect of Trump’s sweeping tariff agenda. But as Maddow revealed, the intellectual foundation for Navarro’s views was alarmingly hollow. In multiple books and memos, he quoted a supposed trade expert named “Ron Vara.” The truth? Ron Vara never existed. He was a fictional character Navarro invented—an anagram of his own last name—used to legitimize arguments that lacked peer-reviewed backing or institutional credibility.<br/><br/>So America’s trade war is, in part, shaped by a dishonest man discovered through an Amazon search, citing a fictional economist he created in his own image. These tariffs aren’t actual strategy. It’s economic theater passed off as “policy” as the world watches the sequel unfold in real time. It’s the most ridiculous thing I’ve ever seen—as China prepares for full-scale economic conflict with the United States, none of this needed to happen.</blockquote> <h3 id="long-term-consequences-and-us-china-relations"><strong>Long Term Consequences and US-China Relations</strong></h3> <p>Next we try to understand the the trajectory of GDP, growth and second and third order effect of this policy (by definition are hard to predict). We also try to summarise the long-term implications for markets and businesses.</p> <p>In a post-coivd world, self-sufficiency is of paramount importance. In recent US has become world’s consumer (woth $20T). The current Trump administration has the mindset that they have the leverage, <strong>if we don’t buy it, they can’t produce it !</strong> <strong>and so we (U.S) are the world’s customer and that the customer is always right.</strong> Its not clear if china can turn all the consumption inward and sell the goods to its own 1 billion.</p> <p>Meanwhile China’s has had its own set of problems. They are facing slowest growth in decades, zero covid policy and crackdown of tech founders etc. lead to a state where its economic growth was 5% in recent years and if trade-wars continue could drop to 2.5% (according to projections by WSJ). Their attractivness as a recepient of forign investment (totalitariism eventually is reconginized loud and clear by momest members of the free countries). They have an aging population and from that prespective it seems they do need exports and can’t absorb everything they create. To stimulate growth govt is investing in tech sector. Other less favorubale options include cocnede defeat to whatever terms trump deamnds. Devalue Yuan by 205 or unleash a fiscal stimular (in the trillions).</p> <p>US is also behind in a number of sectors, like the energy sector while china and india have been building massive electrical generation capacity which will have negative consequeneces as the fight for AI race intensciises (where energy needs are immense).</p> <p>But the debt is unsustainable because the of the large imbalance between a) debtor-borrowers who owe too much debt and are taking on too much debt because they are hooked on debt to finance their excesses (e.g., the United States) and b) lender-creditors (like China) who already hold too much of the debt and are hooked on selling their goods to the borrower-debtors (like the United States) to sustain their economies. There are big pressures for these imbalances to be corrected one way or another and doing so will change the monetary order in major ways. For example, it is obviously incongruous to have both large trade imbalances and large capital imbalances in a deglobalizing world in which the major players can't trust that the other major players won't cut them off from the items they need (which is an American worry) or pay them the money they are owed (which is a Chinese worry).</p> <p>So, the old monetary/economic order in which countries like China manufacture inexpensively, sell to Americans, and acquire American debt assets, and Americans borrow money from countries like China to make those purchases and build up huge debt liabilities will have to change. Clearly, the monetary order will have to change in big disruptive ways to reduce all these imbalances and excesses, and we are in the early part of the process of it changing. There are huge capital market implications to this that have huge economic implications, which I will delve into at another time.</p> <p><strong>similar positions to help you build a list of things that they might do—things like suspending debt service payments to "enemy" countries, establishing capital controls to prevent the free flow of capital out of the country, and imposing special taxes. </strong>Many of these things would’ve been unimaginable not long ago, so <strong>we should also study how these policies work</strong>. The breakdowns in the monetary, political, and geopolitical orders that take the forms of depressions, civil wars, and world wars, that then lead to the new monetary, political orders that govern interactions within countries, and the geopolitical orders that govern interactions between countries until they break down, have all happened repeatedly and are the most important things to understand well.</p> <p>Overall, The multilateral, cooperative world order the U.S. led is being replaced by a unilateral, power-rules approach In this new order, the U.S. is still largest power in the world and is shifting to a unilateral, "America first" approach.</p> <p></p> <p></p> <p></p> <p></p> <p></p> <hr/> <h5 id="references"><strong>References:</strong></h5> <p>I am not an economist. I wrote/compiled the article mainly for myself (using notes from following source).</p> <ol><li><strong>Ray Dalio Article on </strong><a href="https://www.linkedin.com/pulse/dont-make-mistake-thinking-whats-now-happening-mostly-ray-dalio-w8dbe" rel="noopener noreferrer nofollow">Don't Make the Mistake of Thinking That What's Now Happening is Mostly About Tariffs:</a></li><li>All-in Interview with Scott Bessent Tresurer</li><li>All-in interview with Howard Lutnick</li><li><a href="https://chamath.substack.com/p/short-dive-the-trump-administrations?utm_campaign=post&amp;utm_medium=web&amp;triedRedirect=true" rel="noopener noreferrer nofollow"><strong>Short Dive: The Trump Administration's Fiscal Strategy</strong></a></li><li>ChatGPT</li></ol>]]></content><author><name></name></author><summary type="html"><![CDATA[What are the structural problems that the Trump administration is focused on solving? What is this administration’s vision for America? What are the implications for entrepreneurs and investors? The international geopolitical world order is breaking down. Up until now, monetary/economic order dictated that countries like China manufacture inexpensively, sell to Americans, and acquire American debt assets. But we are entering an new era of de-globalization, where rules of monetary, political, and geopolitical orders are being rewritten at an unprecedented rate. This article attempts to understand the biggest disruptions that are likely still ahead. But first we awnser, What are the underlying circumstances that caused US citizens to first elect Donald Trump and then Trump to implement massive tariffs, introduce a DOGE program, adjust taxes and so on., The United States' national debt currently over $34 trillion is the result of a combination of historical, structural, and policy-driven factors that have accumulated over decades. US Public Debt Per Capita is at a current level of 106.11K, up from 104.06K last month and up from 100.41K one year ago. This level of current debt level and the rate at which the government debt is being added to is unsustainable. How Did the US Get here here? Two third of American population is not college educated. Historically upward mobility was possible because there we lots of opportunities to find trade and blue colar jobs. But in recent decades, cheaper Chinese manufacturing - led to deteriorating of American manufacturing, which both hollows out middle class jobs in the U.S. and requires America to import needed items from a country that it is increasingly seeing as an enemy. It also leads to stagnation of wages for the white color jobs. All of this ultimately leads to breakdown of domestic political order due to huge gaps in people's education levels, opportunity levels, productivity levels, income and wealth levels, and values. History also shows that strong autocratic leaders emerge as classic democracy and classic rule of law are removed as barriers to autocratic leadership. At the moment, the U.S. runs a budget deficit when its annual spending exceeds its revenue. These deficits are financed by borrowing (i.e., issuing Treasury securities). Over time, recurring deficits accumulate into large national debt. Even in good economic times, the U.S. often spends more than it collects due to entitlement programs and defense spending. Furthermore, Tax cuts, such as those in 2001 (Bush-era) and 2017 (Trump-era), reduced federal revenue without equivalent spending cuts. In summary, Chronic budget deficits driven by entitlement spending, tax cuts, and military spending have been a driving force with interest rate projected to compound each year until the whole system is no longer sustainable and collapses (like the many empires of the past). The biggest components of U.S. government spending are, Social Security, Medicare, Defense and Interest on debt ! Many of these programs are politically difficult to cut and are growing due to demographic shifts, especially the aging Baby Boomer population (more retirees=more Medicare/Social Security spending). The U.S. has borrowed heavily in response to major events: 2001–2020 wars (Afghanistan, Iraq): trillions spent, mostly borrowed.2008 Global Financial Crisis: stimulus packages and bailouts.2020 COVID-19 pandemic: unprecedented $5+ trillion in relief spending. These crises were partly unavoidable, but they contributed massively to the debt. Interest Rates and Politics: For more than a decade, US has been in a unique global position that has allowed it to borrow more easily than other nations. Additionally, U.S. borrowed at historically low interest rates, which made debt accumulation seem sustainable. Policymakers became less cautious about adding to the debt, as the cost of servicing it was low. There always has been a political reluctance to raise taxes or cut popular programs as the political operation system encourages short-term decision-making. This has now changed: interest payments are the fastest-growing part of the federal budget due to rising rates. The seasoned Politicians can’t seem to do much about it due to political Gridlock and Short-Term Incentives. No long-term fiscal plan: U.S. politicians often avoid unpopular decisions like raising taxes or cutting benefits.Debt ceiling fights don’t reduce debt—they only delay or complicate borrowing.Populist pressure: There’s little political reward for fiscal discipline, especially when voters demand both low taxes and generous public programs. Balancing the Budget: There are two forces to worry about growth and inflation. and therefore, balancing a budget without brining along the congress: Cut Government spending (DOGE)Earn more Understanding the Tarrif Policy: Here is one theory explaining trump Govt. behaviour: During a recent MSNBC segment, Rachel Maddow unpacked the bizarre origin of Donald Trump’s trade policy during his first term. In 2016, seeking to bolster Trump’s economic team, Jared Kushner searched Amazon for books on China. He came across Death by China, co-authored by a little-known UC Irvine professor named Peter Navarro. Impressed by its aggressive tone, Kushner contacted Navarro, who was soon advising Trump directly on trade policy.Navarro became the architect of Trump’s sweeping tariff agenda. But as Maddow revealed, the intellectual foundation for Navarro’s views was alarmingly hollow. In multiple books and memos, he quoted a supposed trade expert named “Ron Vara.” The truth? Ron Vara never existed. He was a fictional character Navarro invented—an anagram of his own last name—used to legitimize arguments that lacked peer-reviewed backing or institutional credibility.So America’s trade war is, in part, shaped by a dishonest man discovered through an Amazon search, citing a fictional economist he created in his own image. These tariffs aren’t actual strategy. It’s economic theater passed off as “policy” as the world watches the sequel unfold in real time. It’s the most ridiculous thing I’ve ever seen—as China prepares for full-scale economic conflict with the United States, none of this needed to happen. Long Term Consequences and US-China Relations Next we try to understand the the trajectory of GDP, growth and second and third order effect of this policy (by definition are hard to predict). We also try to summarise the long-term implications for markets and businesses. In a post-coivd world, self-sufficiency is of paramount importance. In recent US has become world’s consumer (woth $20T). The current Trump administration has the mindset that they have the leverage, if we don’t buy it, they can’t produce it ! and so we (U.S) are the world’s customer and that the customer is always right. Its not clear if china can turn all the consumption inward and sell the goods to its own 1 billion. Meanwhile China’s has had its own set of problems. They are facing slowest growth in decades, zero covid policy and crackdown of tech founders etc. lead to a state where its economic growth was 5% in recent years and if trade-wars continue could drop to 2.5% (according to projections by WSJ). Their attractivness as a recepient of forign investment (totalitariism eventually is reconginized loud and clear by momest members of the free countries). They have an aging population and from that prespective it seems they do need exports and can’t absorb everything they create. To stimulate growth govt is investing in tech sector. Other less favorubale options include cocnede defeat to whatever terms trump deamnds. Devalue Yuan by 205 or unleash a fiscal stimular (in the trillions). US is also behind in a number of sectors, like the energy sector while china and india have been building massive electrical generation capacity which will have negative consequeneces as the fight for AI race intensciises (where energy needs are immense). But the debt is unsustainable because the of the large imbalance between a) debtor-borrowers who owe too much debt and are taking on too much debt because they are hooked on debt to finance their excesses (e.g., the United States) and b) lender-creditors (like China) who already hold too much of the debt and are hooked on selling their goods to the borrower-debtors (like the United States) to sustain their economies. There are big pressures for these imbalances to be corrected one way or another and doing so will change the monetary order in major ways. For example, it is obviously incongruous to have both large trade imbalances and large capital imbalances in a deglobalizing world in which the major players can't trust that the other major players won't cut them off from the items they need (which is an American worry) or pay them the money they are owed (which is a Chinese worry). So, the old monetary/economic order in which countries like China manufacture inexpensively, sell to Americans, and acquire American debt assets, and Americans borrow money from countries like China to make those purchases and build up huge debt liabilities will have to change. Clearly, the monetary order will have to change in big disruptive ways to reduce all these imbalances and excesses, and we are in the early part of the process of it changing. There are huge capital market implications to this that have huge economic implications, which I will delve into at another time. similar positions to help you build a list of things that they might do—things like suspending debt service payments to "enemy" countries, establishing capital controls to prevent the free flow of capital out of the country, and imposing special taxes. Many of these things would’ve been unimaginable not long ago, so we should also study how these policies work. The breakdowns in the monetary, political, and geopolitical orders that take the forms of depressions, civil wars, and world wars, that then lead to the new monetary, political orders that govern interactions within countries, and the geopolitical orders that govern interactions between countries until they break down, have all happened repeatedly and are the most important things to understand well. Overall, The multilateral, cooperative world order the U.S. led is being replaced by a unilateral, power-rules approach In this new order, the U.S. is still largest power in the world and is shifting to a unilateral, "America first" approach. References: I am not an economist. I wrote/compiled the article mainly for myself (using notes from following source). Ray Dalio Article on Don't Make the Mistake of Thinking That What's Now Happening is Mostly About Tariffs:All-in Interview with Scott Bessent TresurerAll-in interview with Howard LutnickShort Dive: The Trump Administration's Fiscal StrategyChatGPT]]></summary></entry><entry><title type="html">Accelerating Scientific Discovery with AI</title><link href="https://asjad99.github.io/blog/2026/04/accelerating-scientific-discovery-with-ai/" rel="alternate" type="text/html" title="Accelerating Scientific Discovery with AI"/><published>2026-04-11T10:25:40+00:00</published><updated>2026-04-11T10:25:40+00:00</updated><id>https://asjad99.github.io/blog/2026/04/accelerating-scientific-discovery-with-ai</id><content type="html" xml:base="https://asjad99.github.io/blog/2026/04/accelerating-scientific-discovery-with-ai/"><![CDATA[<p><em>The most beautiful thing you can experience is mysterious&nbsp;- Einstein </em></p> <p>One area that caught my attention was Compute driven scientific discovery. So i have been trying to understand from first principles what role computing/AI will play in scientific discovery in coming years. Starting from a blank canvas, i thought it ’d be best to write down some stuff and possibly see if a picture emerges and if we can also find some common ground for projects and develop a vision statement for atreides research lab. Here are some long, random and incomplete thoughts starting from very basics:</p> <p></p> <blockquote>THE SCIENTIFIC method was perhaps the single most important development in modern history. It established a way to validate truth at a time when misinformation was the norm, allowing natural philosophers to navigate the unknown. From predicting the motions of the planets to discovering the principles of electricity, scientists have honed the ability to distil facts about the universe by generating theories, then using experimentation to qualify those theories. Looking at how far civilisation has come since the Enlightenment, one can’t help being awestruck by all that humanity has achieved using this approach. I believe artificial intelligence (AI) could usher in a new renaissance of discovery, acting as a multiplier for human ingenuity, opening up entirely new areas of inquiry and spurring humanity to realise its full potential. The promise of AI is that it could serve as an extension of our minds and become a meta-solution. In the same way that the telescope revealed the planetary dynamics that inspired new physics, insights from AI could help scientists solve some of the complex challenges facing society today—from superbugs to climate change to inequality. My hope is to build smarter tools that expand humans’ capacity to identify the root causes and potential solutions to core scientific problems<a href="https://worldin.economist.com/article/17385/edition2020demis-hassabis-predicts-ai-will-supercharge-science?ref=asjadkhan.ghost.io">.&nbsp;<strong>Demis Hassabis on AI's potential</strong></a></blockquote> <p></p> <p>Folks with a CS background are trained in algorithmic thinking where we learn to thinking about solving problems using computation. Algorithms are the fundamental unit of programming and computer science. But increasingly in terms of applications they go beyond software development as well. Avi Higderson says <em>“Algorithms are a common language for nature, human, and computer.”</em> Even though the field of CS is new and increasingly we are starting to realise that algorithmic thinking is a universal framework that can be applied to hard sciences as well. Robert Sedgewick who teaches algorithms at Princeton says computational models are replacing math models in scientific inquiry/discovery:</p> <figure class="kg-card kg-image-card"><img src="/assets/img/ghost/d3034d98ff-image.png" class="kg-image" alt="" loading="lazy" width="1204" height="396" srcset="/assets/img/ghost/4ca829b40d-image.png 600w, /assets/img/ghost/29f15efe3e-image.png 1000w, /assets/img/ghost/d3034d98ff-image.png 1204w" sizes="(min-width: 720px) 720px"/></figure> <p>This means algorithms are a common language for understanding nature where we can simulate the given phenomena in order to better understand it. This attitude however, does have its limitations:</p> <p>Data centric computing and Data Science computer science has evolved as a discipline. Its interesting to see how MIT’s intro to computation course has changed its structure overtime with emphasis on technical topics for Data analysis as well. AI Software is replacing traditional software Some dub the phenomena as software 2.0. I like the term data-centric computing for explaining this phenomena.</p> <figure class="kg-card kg-image-card"><img src="/assets/img/ghost/9053a84c9a-image-1.png" class="kg-image" alt="" loading="lazy" width="1140" height="440" srcset="/assets/img/ghost/b74591aac6-image-1.png 600w, /assets/img/ghost/c437f34478-image-1.png 1000w, /assets/img/ghost/9053a84c9a-image-1.png 1140w" sizes="(min-width: 720px) 720px"/></figure> <p>Data science is the process of formulating a quantitative question that can be answered with data, collecting and cleaning the data, analyzing the data, and communicating the answer to the question to a relevant audience. If we do want a concise definition, the following seems to be reasonable: Data science is the application of computational and statistical techniques to address or gain insight into some problem in the real world.</p> <p>The key phrases of importance here are “computational” (data science typically involves some sort of algorithmic methods written in code), “statistical” (statistical inference lets us build the predictions that we make), and “real world” (we are talking about deriving insight not into some artificial process, but into some “truth” in the real world).</p> <p>Another way of looking at it, in some sense, is that data science is simply the union of the various techniques that are required to accomplish the above. That is, something like: </p> <p><strong>Data science</strong> = statistics + data collection + data preprocessing + machine learning + visualisation + business insights + scientific hypotheses + big data + (etc) </p> <p>This definition is also useful, mainly because it emphasizes that all these areas are crucial to obtaining the goals of data science. In fact, in some sense data science is best defined in terms of what it is not, namely, that is it not (just) any one of these subjects above. </p> <p><strong>Data centric</strong> = data science + data structures </p> <p><strong>Data science</strong> = statistics + data collection + data preprocessing + machine learning + visualisation + business insights + scientific hypotheses + big data + (etc) </p> <p>Data centric computing is characterised by programs interrogate data: that is, programs are tools for answering questions. <a href="https://dl.acm.org/doi/abs/10.1145/3408877.3432457">read more</a></p> <p><strong>AI as an enabler of scientific discovery - Case Studies</strong></p> <p>In the mid-twentieth century, Margaret Oakley Dayhoff pioneered the analysis of protein sequencing data, a forerunner of genome sequencing, leading early research that used computers to analyse patterns in the sequences. The expression ‘artificial intelligence’ today is therefore an umbrella term. It refers to a suite of technologies that can perform complex tasks when acting in conditions of uncertainty, including visual perception, speech recognition, natural language processing, reasoning, learning from data, and a range of optimisation problems.</p> <p><strong>Using genomic data to predict protein structures:</strong> Understanding a protein’s shape is key to understanding the role it plays in the body. By predicting these shapes, scientists can identify proteins that play a role in diseases, improving diagnosis and helping develop new treatments. The process of determining protein structures is both technically difficult and labour-intensive, yielding approximately 100,000 known structures to date5. While advances in genetics in recent decades have provided rich datasets of DNA sequences, determining the shape of a protein from its corresponding genetic sequence – the protein-folding challenge – is a complex task. To help understand this process, researchers are developing machine learning approaches that can predict the threedimensional structure of proteins from DNA sequences. The AlphaFold project at DeepMind, for example, has created a deep neural network that predicts the distances between pairs of amino acids and the angles between their bonds, and in so doing produces a highly-accurate prediction of an overall protein structure.</p> <p><strong>Understanding complex organic chemistry </strong>The goal of this pilot project between the John Innes Centre and The Alan Turing Institute is to investigate possibilities for machine learning in modelling and predicting the process of triterpene biosynthesis in plants. Triterpenes are complex molecules which form a large and important class of plant natural products, with diverse commercial applications across the health, agriculture and industrial sectors. The triterpenes are all synthesized from a single common substrate which can then be further modified by tailoring enzymes to give over 20,000 structurally diverse triterpenes. Recent machine learning models have shown promise at predicting the outcomes of organic chemical reactions. Successful prediction based on sequence will require both a deep understanding of the biosynthetic pathways that produce triterpenes, as well as novel machine learning methodology</p> <p>• Finding patterns in astronomical data: Driving scientific discovery from particle physics experiments and large scale astronomical data • Understanding the effects of climate change on cities and regions: Satellite imaging to support conservation</p> <p><em>The ‘traditional’ way to apply data science methods is to start from a large data set, and then apply machine learning methods to try to discover patterns that are hidden in the data – without taking into account anything about where the data came from, or current knowledge of the system. But might it be possible to incorporate existing scientific knowledge (for example, in the form of a statistical ‘prior’) so that the discovery process is constrained, in order to produce results which respect what researchers already know about the system. For example, if trying to detect the 3D shape of a protein from image data, could chemical knowledge of how proteins fold be incorporated in the analysis, in order to guide the search? the goal of scientific discovery is to understand. Researchers want to know not just what the answer is but why. Are there ways of using AI algorithms that will provide such explanations? In what ways might AI-enabled analysis and hypothesis-led research sit alongside each other in future? How might people work with AI to solve scientific mysteries in the years to come? Is it possible that one day, computational methods will not only discover patterns and unusual events in data, but have enough domain knowledge built in that they can themselves make new scientific breakthroughs? Could they come up with new theories that revolutionise our understanding, and devise novel experiments to test them out? Could they even decide for themselves what the worthwhile scientific questions are? And worthwhile to whom? </em></p> <p>Scientific Method is about building Error Correction Systems - ones that become better over time. </p>]]></content><author><name></name></author><summary type="html"><![CDATA[The most beautiful thing you can experience is mysterious&nbsp;- Einstein One area that caught my attention was Compute driven scientific discovery. So i have been trying to understand from first principles what role computing/AI will play in scientific discovery in coming years. Starting from a blank canvas, i thought it ’d be best to write down some stuff and possibly see if a picture emerges and if we can also find some common ground for projects and develop a vision statement for atreides research lab. Here are some long, random and incomplete thoughts starting from very basics: THE SCIENTIFIC method was perhaps the single most important development in modern history. It established a way to validate truth at a time when misinformation was the norm, allowing natural philosophers to navigate the unknown. From predicting the motions of the planets to discovering the principles of electricity, scientists have honed the ability to distil facts about the universe by generating theories, then using experimentation to qualify those theories. Looking at how far civilisation has come since the Enlightenment, one can’t help being awestruck by all that humanity has achieved using this approach. I believe artificial intelligence (AI) could usher in a new renaissance of discovery, acting as a multiplier for human ingenuity, opening up entirely new areas of inquiry and spurring humanity to realise its full potential. The promise of AI is that it could serve as an extension of our minds and become a meta-solution. In the same way that the telescope revealed the planetary dynamics that inspired new physics, insights from AI could help scientists solve some of the complex challenges facing society today—from superbugs to climate change to inequality. My hope is to build smarter tools that expand humans’ capacity to identify the root causes and potential solutions to core scientific problems.&nbsp;Demis Hassabis on AI's potential Folks with a CS background are trained in algorithmic thinking where we learn to thinking about solving problems using computation. Algorithms are the fundamental unit of programming and computer science. But increasingly in terms of applications they go beyond software development as well. Avi Higderson says “Algorithms are a common language for nature, human, and computer.” Even though the field of CS is new and increasingly we are starting to realise that algorithmic thinking is a universal framework that can be applied to hard sciences as well. Robert Sedgewick who teaches algorithms at Princeton says computational models are replacing math models in scientific inquiry/discovery: This means algorithms are a common language for understanding nature where we can simulate the given phenomena in order to better understand it. This attitude however, does have its limitations: Data centric computing and Data Science computer science has evolved as a discipline. Its interesting to see how MIT’s intro to computation course has changed its structure overtime with emphasis on technical topics for Data analysis as well. AI Software is replacing traditional software Some dub the phenomena as software 2.0. I like the term data-centric computing for explaining this phenomena. Data science is the process of formulating a quantitative question that can be answered with data, collecting and cleaning the data, analyzing the data, and communicating the answer to the question to a relevant audience. If we do want a concise definition, the following seems to be reasonable: Data science is the application of computational and statistical techniques to address or gain insight into some problem in the real world. The key phrases of importance here are “computational” (data science typically involves some sort of algorithmic methods written in code), “statistical” (statistical inference lets us build the predictions that we make), and “real world” (we are talking about deriving insight not into some artificial process, but into some “truth” in the real world). Another way of looking at it, in some sense, is that data science is simply the union of the various techniques that are required to accomplish the above. That is, something like: Data science=statistics + data collection + data preprocessing + machine learning + visualisation + business insights + scientific hypotheses + big data + (etc) This definition is also useful, mainly because it emphasizes that all these areas are crucial to obtaining the goals of data science. In fact, in some sense data science is best defined in terms of what it is not, namely, that is it not (just) any one of these subjects above. Data centric=data science + data structures Data science=statistics + data collection + data preprocessing + machine learning + visualisation + business insights + scientific hypotheses + big data + (etc) Data centric computing is characterised by programs interrogate data: that is, programs are tools for answering questions. read more AI as an enabler of scientific discovery - Case Studies In the mid-twentieth century, Margaret Oakley Dayhoff pioneered the analysis of protein sequencing data, a forerunner of genome sequencing, leading early research that used computers to analyse patterns in the sequences. The expression ‘artificial intelligence’ today is therefore an umbrella term. It refers to a suite of technologies that can perform complex tasks when acting in conditions of uncertainty, including visual perception, speech recognition, natural language processing, reasoning, learning from data, and a range of optimisation problems. Using genomic data to predict protein structures: Understanding a protein’s shape is key to understanding the role it plays in the body. By predicting these shapes, scientists can identify proteins that play a role in diseases, improving diagnosis and helping develop new treatments. The process of determining protein structures is both technically difficult and labour-intensive, yielding approximately 100,000 known structures to date5. While advances in genetics in recent decades have provided rich datasets of DNA sequences, determining the shape of a protein from its corresponding genetic sequence – the protein-folding challenge – is a complex task. To help understand this process, researchers are developing machine learning approaches that can predict the threedimensional structure of proteins from DNA sequences. The AlphaFold project at DeepMind, for example, has created a deep neural network that predicts the distances between pairs of amino acids and the angles between their bonds, and in so doing produces a highly-accurate prediction of an overall protein structure. Understanding complex organic chemistry The goal of this pilot project between the John Innes Centre and The Alan Turing Institute is to investigate possibilities for machine learning in modelling and predicting the process of triterpene biosynthesis in plants. Triterpenes are complex molecules which form a large and important class of plant natural products, with diverse commercial applications across the health, agriculture and industrial sectors. The triterpenes are all synthesized from a single common substrate which can then be further modified by tailoring enzymes to give over 20,000 structurally diverse triterpenes. Recent machine learning models have shown promise at predicting the outcomes of organic chemical reactions. Successful prediction based on sequence will require both a deep understanding of the biosynthetic pathways that produce triterpenes, as well as novel machine learning methodology • Finding patterns in astronomical data: Driving scientific discovery from particle physics experiments and large scale astronomical data • Understanding the effects of climate change on cities and regions: Satellite imaging to support conservation The ‘traditional’ way to apply data science methods is to start from a large data set, and then apply machine learning methods to try to discover patterns that are hidden in the data – without taking into account anything about where the data came from, or current knowledge of the system. But might it be possible to incorporate existing scientific knowledge (for example, in the form of a statistical ‘prior’) so that the discovery process is constrained, in order to produce results which respect what researchers already know about the system. For example, if trying to detect the 3D shape of a protein from image data, could chemical knowledge of how proteins fold be incorporated in the analysis, in order to guide the search? the goal of scientific discovery is to understand. Researchers want to know not just what the answer is but why. Are there ways of using AI algorithms that will provide such explanations? In what ways might AI-enabled analysis and hypothesis-led research sit alongside each other in future? How might people work with AI to solve scientific mysteries in the years to come? Is it possible that one day, computational methods will not only discover patterns and unusual events in data, but have enough domain knowledge built in that they can themselves make new scientific breakthroughs? Could they come up with new theories that revolutionise our understanding, and devise novel experiments to test them out? Could they even decide for themselves what the worthwhile scientific questions are? And worthwhile to whom? Scientific Method is about building Error Correction Systems - ones that become better over time.]]></summary></entry><entry><title type="html">Apple’s Approach to Design thinking</title><link href="https://asjad99.github.io/blog/2024/12/design-thinking/" rel="alternate" type="text/html" title="Apple’s Approach to Design thinking"/><published>2024-12-21T09:01:37+00:00</published><updated>2024-12-21T09:01:37+00:00</updated><id>https://asjad99.github.io/blog/2024/12/design-thinking</id><content type="html" xml:base="https://asjad99.github.io/blog/2024/12/design-thinking/"><![CDATA[<p>Design thinking has become something of a buzzword in recent years. But behind the buzz is a genuinely transformative way of approaching problems one that centers on people. This framework, known as <strong>Human-Centered Design (HCD)</strong>, is about more than just solving problems; it’s about solving the right problems. And the key lies in deeply understanding human needs, cultures, and contexts.</p> <p>I recently attended a Masterclass on design thinking, and it left me reflecting on how powerful this approach can be—not just for businesses but for anyone grappling with complex challenges. The insights I walked away with underscored how HCD has evolved, why it matters, and how its principles can be applied effectively. But before we jump into that - let me share Steve Job's thought how merging humanities with sciences was the key to apple's success. </p> <h3 id="steve-jobs-on-apple-experiences">Steve jobs on Apple experiences: </h3> <blockquote><br/> "There were a lot of people at Apple that just didn't get it. We fought tooth and nail with a variety of people there who thought the whole concept of a graphical user interface was crazy ... on the grounds that it couldn't be done, or on the grounds that real computer users didn't need menus in plain English, and real computer users didn't care about putting nice little pictures on the screen. But fortunately, I was the largest stockholder and the chairman of the company, so I won. Apple was a corporation, we were very conscious of that. We were driven to make money. I would say that Apple was a corporate lifestyle, but it had a few big differences from other corporate lifestyles I'd seen. The first one was a real belief that there wasn't a hierarchy of ideas that mapped into the hierarchy of the organization. In other words: great ideas could come from anywhere. Apple was a very bottom-up company when it came to a lot of its great ideas. We hired truly great people and gave them the room to do great work. A lot of companies — I know it sounds crazy — but a lot of companies don't do that. They hire people to tell them what to do. We hire people to tell <em>us</em> what to do. We figure we're paying them all this money; their job is to figure out what to do and tell us. That led to a very different corporate culture, and one that's really much more collegial than hierarchical.<strong> I think our major contribution [to computing] was in bringing a liberal arts point of view to the use of computers.</strong> If you really look at the ease of use of the Macintosh, the driving motivation behind that was to bring not only ease of use to people — so that many, many more people could use computers for nontraditional things at that time — but it was to bring beautiful fonts and typography to people, it was to bring graphics to people ... so that they could see beautiful photographs, or pictures, or artwork, et cetera ... to help them communicate. ... Our goal was to bring a liberal arts perspective and a liberal arts audience to what had traditionally been a very geeky technology and a very geeky audience. In my perspective ... science and<strong> computer science <em>is</em> a liberal art, </strong>it's something everyone should know how to use, at least, and harness in their life. It's not something that should be relegated to 5 percent of the population over in the corner. It's something that everybody should be exposed to and everyone should have mastery of to some extent, and that's how we viewed computation and these computation devices.”<br/><em>“I think <strong>great artists and great engineers are similar</strong>, in that they both have a desire to express themselves. In fact some of the best people working on the original Mac were poets and musicians on the side.” </em></blockquote> <p></p> <figure class="kg-card kg-image-card"><img src="/assets/img/ghost/c4450c3bf1-image-1.png" class="kg-image" alt="" loading="lazy" width="1290" height="724" srcset="/assets/img/ghost/dc027e15bf-image-1.png 600w, /assets/img/ghost/e11ced0346-image-1.png 1000w, /assets/img/ghost/c4450c3bf1-image-1.png 1290w" sizes="(min-width: 720px) 720px"/></figure> <figure class="kg-card kg-image-card"><img src="/assets/img/ghost/2e0925445d-image.png" class="kg-image" alt="" loading="lazy" width="1059" height="596" srcset="/assets/img/ghost/d948dcf9f6-image.png 600w, /assets/img/ghost/34742036c8-image.png 1000w, /assets/img/ghost/2e0925445d-image.png 1059w" sizes="(min-width: 720px) 720px"/></figure> <p>HCD isn’t a new idea. Its roots stretch back to the early 20th century, when industrial designers began thinking about ergonomics and usability. But it wasn’t until the cognitive revolution of the 1960s that the field really began to take shape. The rise of personal computing in the 1970s and 80s brought usability to the forefront, and by the early 2000s, companies like IDEO had popularized design thinking as a framework for tackling not just digital challenges but broader organizational ones.</p> <p>Today, the scope of HCD has expanded further. Inclusive design is now a major focus, emphasizing the need to consider diverse users and contexts. It’s no longer just about making things work—it’s about making them work for everyone. One of the most compelling ideas from the Masterclass was that HCD isn’t just about designing solutions. It’s about designing solutions that matter. To do this, you have to start by understanding the people you’re designing for—their needs, limitations, and goals.</p> <p>This focus on people has a real impact. A striking statistic shared during the session highlighted that companies integrating design thinking into their strategy can outperform industry peers by as much as <strong>228%</strong>. But the benefits go beyond revenue. Seventy-one percent of companies reported that design thinking improved their working culture, leading to greater productivity and employee engagement.</p> <p>At the heart of design thinking lies the <strong>Double Diamond Framework</strong>, a simple yet profound way of visualizing the design process. It’s divided into four phases, alternating between divergent and convergent thinking:</p> <ol><li><strong>Discover</strong>: This is where you explore the problem space, gathering insights into users’ needs and challenges. It’s about asking questions, not rushing to answers.</li><li><strong>Define</strong>: Once you’ve gathered enough data, the focus shifts to clarity. What’s the real problem? This step sets the stage for everything that follows.</li><li><strong>Develop</strong>: With a clear problem in mind, it’s time to brainstorm solutions. This is the phase for creativity, experimentation, and pushing boundaries.</li><li><strong>Deliver</strong>: Finally, you narrow down and refine your solutions, creating prototypes, testing them, and iterating until you arrive at something ready to launch.</li></ol> <p>The beauty of this framework lies in its structure. It provides a clear path forward, but it’s flexible enough to adapt to the messy, iterative nature of real-world problem-solving. The Masterclass also introduced a range of tools that can be used throughout the design process. Here are a few that stood out:</p> <ul><li><strong>Empathy Mapping</strong>: This tool helps teams step into users’ shoes by exploring their thoughts, feelings, and behaviors.</li><li><strong>Journey Mapping</strong>: By visualizing the user’s experience across different touchpoints, you can identify pain points and opportunities for improvement.</li><li><strong>Rapid Prototyping</strong>: Quickly building and testing ideas allows for fast feedback and iteration.</li><li><strong>User Interviews</strong>: Structured conversations with users provide invaluable insights into their needs and preferences.</li></ul> <p>Each of these tools reinforces the idea that good design starts with good listening.</p> <p>One of the most important lessons from the Masterclass was the distinction between <strong>problem explorers</strong> and <strong>problem solvers</strong>. Too often, we rush to solutions without fully understanding the problem. This leads to what’s known as “designer myopia,” where solutions may impress peers but fail to meet users’ actual needs.</p> <p>The design thinking framework forces you to slow down and explore. It emphasizes that the most innovative solutions often emerge from a deep understanding of the problem space. And that understanding doesn’t come from sitting in a conference room—it comes from engaging with real people in real contexts.</p> <p>Ultimately, HCD isn’t just about creating functional solutions. It’s about creating solutions that resonate—solutions that are meaningful, sustainable, and deeply human. The structured yet flexible nature of the Double Diamond Framework makes it an invaluable tool for navigating uncertainty, exploring diverse ideas, and delivering outcomes that matter.</p> <p>The real power of design thinking lies in its ability to align creativity with purpose. By centering on human needs and encouraging collaboration across disciplines, it transforms not just what we create but how we create. And in doing so, it opens the door to solutions that truly make a difference. Design thinking isn’t just a process; it’s a mindset. It’s about curiosity, empathy, and the willingness to embrace complexity. The lessons from the Masterclass reinforced the idea that by staying grounded in human-centered principles, we can tackle even the most challenging problems with confidence and creativity.</p> <p></p>]]></content><author><name></name></author><summary type="html"><![CDATA[Design thinking has become something of a buzzword in recent years. But behind the buzz is a genuinely transformative way of approaching problems one that centers on people. This framework, known as Human-Centered Design (HCD), is about more than just solving problems; it’s about solving the right problems. And the key lies in deeply understanding human needs, cultures, and contexts. I recently attended a Masterclass on design thinking, and it left me reflecting on how powerful this approach can be—not just for businesses but for anyone grappling with complex challenges. The insights I walked away with underscored how HCD has evolved, why it matters, and how its principles can be applied effectively. But before we jump into that - let me share Steve Job's thought how merging humanities with sciences was the key to apple's success. Steve jobs on Apple experiences: "There were a lot of people at Apple that just didn't get it. We fought tooth and nail with a variety of people there who thought the whole concept of a graphical user interface was crazy ... on the grounds that it couldn't be done, or on the grounds that real computer users didn't need menus in plain English, and real computer users didn't care about putting nice little pictures on the screen. But fortunately, I was the largest stockholder and the chairman of the company, so I won. Apple was a corporation, we were very conscious of that. We were driven to make money. I would say that Apple was a corporate lifestyle, but it had a few big differences from other corporate lifestyles I'd seen. The first one was a real belief that there wasn't a hierarchy of ideas that mapped into the hierarchy of the organization. In other words: great ideas could come from anywhere. Apple was a very bottom-up company when it came to a lot of its great ideas. We hired truly great people and gave them the room to do great work. A lot of companies — I know it sounds crazy — but a lot of companies don't do that. They hire people to tell them what to do. We hire people to tell us what to do. We figure we're paying them all this money; their job is to figure out what to do and tell us. That led to a very different corporate culture, and one that's really much more collegial than hierarchical. I think our major contribution [to computing] was in bringing a liberal arts point of view to the use of computers. If you really look at the ease of use of the Macintosh, the driving motivation behind that was to bring not only ease of use to people — so that many, many more people could use computers for nontraditional things at that time — but it was to bring beautiful fonts and typography to people, it was to bring graphics to people ... so that they could see beautiful photographs, or pictures, or artwork, et cetera ... to help them communicate. ... Our goal was to bring a liberal arts perspective and a liberal arts audience to what had traditionally been a very geeky technology and a very geeky audience. In my perspective ... science and computer science is a liberal art, it's something everyone should know how to use, at least, and harness in their life. It's not something that should be relegated to 5 percent of the population over in the corner. It's something that everybody should be exposed to and everyone should have mastery of to some extent, and that's how we viewed computation and these computation devices.”“I think great artists and great engineers are similar, in that they both have a desire to express themselves. In fact some of the best people working on the original Mac were poets and musicians on the side.” HCD isn’t a new idea. Its roots stretch back to the early 20th century, when industrial designers began thinking about ergonomics and usability. But it wasn’t until the cognitive revolution of the 1960s that the field really began to take shape. The rise of personal computing in the 1970s and 80s brought usability to the forefront, and by the early 2000s, companies like IDEO had popularized design thinking as a framework for tackling not just digital challenges but broader organizational ones. Today, the scope of HCD has expanded further. Inclusive design is now a major focus, emphasizing the need to consider diverse users and contexts. It’s no longer just about making things work—it’s about making them work for everyone. One of the most compelling ideas from the Masterclass was that HCD isn’t just about designing solutions. It’s about designing solutions that matter. To do this, you have to start by understanding the people you’re designing for—their needs, limitations, and goals. This focus on people has a real impact. A striking statistic shared during the session highlighted that companies integrating design thinking into their strategy can outperform industry peers by as much as 228%. But the benefits go beyond revenue. Seventy-one percent of companies reported that design thinking improved their working culture, leading to greater productivity and employee engagement. At the heart of design thinking lies the Double Diamond Framework, a simple yet profound way of visualizing the design process. It’s divided into four phases, alternating between divergent and convergent thinking: Discover: This is where you explore the problem space, gathering insights into users’ needs and challenges. It’s about asking questions, not rushing to answers.Define: Once you’ve gathered enough data, the focus shifts to clarity. What’s the real problem? This step sets the stage for everything that follows.Develop: With a clear problem in mind, it’s time to brainstorm solutions. This is the phase for creativity, experimentation, and pushing boundaries.Deliver: Finally, you narrow down and refine your solutions, creating prototypes, testing them, and iterating until you arrive at something ready to launch. The beauty of this framework lies in its structure. It provides a clear path forward, but it’s flexible enough to adapt to the messy, iterative nature of real-world problem-solving. The Masterclass also introduced a range of tools that can be used throughout the design process. Here are a few that stood out: Empathy Mapping: This tool helps teams step into users’ shoes by exploring their thoughts, feelings, and behaviors.Journey Mapping: By visualizing the user’s experience across different touchpoints, you can identify pain points and opportunities for improvement.Rapid Prototyping: Quickly building and testing ideas allows for fast feedback and iteration.User Interviews: Structured conversations with users provide invaluable insights into their needs and preferences. Each of these tools reinforces the idea that good design starts with good listening. One of the most important lessons from the Masterclass was the distinction between problem explorers and problem solvers. Too often, we rush to solutions without fully understanding the problem. This leads to what’s known as “designer myopia,” where solutions may impress peers but fail to meet users’ actual needs. The design thinking framework forces you to slow down and explore. It emphasizes that the most innovative solutions often emerge from a deep understanding of the problem space. And that understanding doesn’t come from sitting in a conference room—it comes from engaging with real people in real contexts. Ultimately, HCD isn’t just about creating functional solutions. It’s about creating solutions that resonate—solutions that are meaningful, sustainable, and deeply human. The structured yet flexible nature of the Double Diamond Framework makes it an invaluable tool for navigating uncertainty, exploring diverse ideas, and delivering outcomes that matter. The real power of design thinking lies in its ability to align creativity with purpose. By centering on human needs and encouraging collaboration across disciplines, it transforms not just what we create but how we create. And in doing so, it opens the door to solutions that truly make a difference. Design thinking isn’t just a process; it’s a mindset. It’s about curiosity, empathy, and the willingness to embrace complexity. The lessons from the Masterclass reinforced the idea that by staying grounded in human-centered principles, we can tackle even the most challenging problems with confidence and creativity.]]></summary></entry><entry><title type="html">The Cost of Chasing Unicorns: Lessons from the Product Bubble</title><link href="https://asjad99.github.io/blog/2024/10/the-cost-of-chasing-unicorns-lessons-from-the-product-bubble/" rel="alternate" type="text/html" title="The Cost of Chasing Unicorns: Lessons from the Product Bubble"/><published>2024-10-15T21:50:12+00:00</published><updated>2024-10-15T21:50:12+00:00</updated><id>https://asjad99.github.io/blog/2024/10/the-cost-of-chasing-unicorns-lessons-from-the-product-bubble</id><content type="html" xml:base="https://asjad99.github.io/blog/2024/10/the-cost-of-chasing-unicorns-lessons-from-the-product-bubble/"><![CDATA[<p>The venture capital (VC) world has always been a bit like a poker game, with big bets placed on uncertain outcomes. In the recent "Product Bubble," this game reached a fever pitch. VC investments soared from $45 billion to an astonishing $620 billion, chasing an elusive prey: unicorns—startups valued at over $1 billion. This surge brought big wins for a few, but also costs that rippled across industries, communities, and even society at large.</p> <p>But what was really happening beneath the surface? To understand, we need to trace the incentives that shaped this frenzy.</p> <p>Venture capital operates on a principle borrowed from poker: creating “outs.” The idea is to spread bets across high-risk startups, hoping a few hit it big. It’s not about avoiding failure; failure is baked into the model. The aim is to make sure the winners pay for all the losers—and then some.</p> <p>But in the Product Bubble, the stakes were raised. The flood of capital created immense pressure to deliver outsized returns. And delivering those returns required chasing unicorns. Unicorns are extraordinary by definition: billion-dollar companies born from ideas that often seemed ordinary at first glance. The problem was that creating them on such a large scale wasn’t sustainable.</p> <p>At its peak, VC firms were deploying $300 billion annually into startups. The math behind their portfolios was brutal:</p> <ul><li>10% of investments would be home runs, generating most of the returns.</li><li>30% would break even or deliver modest returns.</li><li>60% would fail outright.</li></ul> <p>To meet these odds, VCs had to make a lot of bets. And when there’s too much money chasing too few good ideas, standards inevitably slip. Marginal ideas—the ones that wouldn’t have made the cut in leaner times—got funded. Suddenly, the market was full of startups promising billion-dollar visions, many of which had no business being there.</p> <p>To understand how this played out, consider the kinds of businesses VCs funded:</p> <ul><li>A SaaS analytics company with $70 million in annual recurring revenue? That’s plausible.</li><li>A government contractor with $80 million in EBITDA? Also plausible, with the right patience.</li><li>A social network with 5 million daily users? Feasible with viral growth.</li></ul> <p>But during the Product Bubble, it wasn’t enough to build a solid business. To attract VC attention, startups had to promise hyper-growth. That meant taking risks, stretching visions, and often pushing boundaries in ways that led to unintended consequences.</p> <p>The defining feature of this era was hyper-scale disruption. Startups weren’t just trying to compete; they were trying to upend entire industries. Uber and Airbnb are classic examples. Both started with seemingly innocent ideas—ride-sharing and room-sharing—but grew into juggernauts that reshaped cities, industries, and livelihoods.</p> <p>This kind of disruption had a dark side. Regulations were treated as obstacles to be skirted. Workers became disposable cogs. Quality and sustainability often took a back seat to growth. Consider the gig economy companies that flooded markets with subsidized pricing to undercut competitors. It worked—for a while. But when funding dried up, these models collapsed, leaving industries in disarray and consumers with fewer choices.</p> <p>The Product Bubble is filled with stories of founders whose good intentions led to unintended consequences:</p> <ul><li>A food delivery service grew into a platform that strained restaurants and exploited drivers.</li><li>A room-sharing idea morphed into Airbnb, accused of driving up rents and displacing residents.</li><li>A social app for friends turned into Facebook, implicated in misinformation and polarization.</li></ul> <p>These founders didn’t set out to harm anyone. They were chasing growth. But in doing so, they often overlooked the broader impacts of their decisions.</p> <p>By the end of the bubble, the unicorns of the era—Uber, Airbnb, and others—achieved valuations of $100 billion or more. But their success came at a cost: skyrocketing rents, increased congestion, and precarious working conditions. The relentless pursuit of scale had outstripped the ability to manage it responsibly.</p> <p>Why did this happen? Because the system was optimized for growth at all costs. Startups were incentivized to over-promise and under-deliver, cutting corners and sometimes ignoring regulations entirely. It worked, for a time. But now the bill is due.</p> <p>The Product Bubble has burst, thanks in part to rising interest rates and a cooling economy. This is a moment for everyone—founders, investors, and communities—to reflect. Does every idea need to become a billion-dollar company? Is chasing the biggest possible returns always the best approach? The answers seem increasingly clear: no.</p> <p>The future of venture capital doesn’t have to look like the Product Bubble. The focus can shift toward sustainable growth, responsible innovation, and long-term value. Founders can be encouraged to build resilient businesses that solve real problems, rather than chasing inflated valuations. The Product Bubble was a lesson in excess. It showed what happens when growth becomes an end in itself. But as the dust settles, there’s an opportunity to chart a better course. Future startups can be built on foundations of resilience, responsibility, and true value. The next era of innovation doesn’t have to be a bubble. It can be something better.</p>]]></content><author><name></name></author><summary type="html"><![CDATA[The venture capital (VC) world has always been a bit like a poker game, with big bets placed on uncertain outcomes. In the recent "Product Bubble," this game reached a fever pitch. VC investments soared from $45 billion to an astonishing $620 billion, chasing an elusive prey: unicorns—startups valued at over $1 billion. This surge brought big wins for a few, but also costs that rippled across industries, communities, and even society at large. But what was really happening beneath the surface? To understand, we need to trace the incentives that shaped this frenzy. Venture capital operates on a principle borrowed from poker: creating “outs.” The idea is to spread bets across high-risk startups, hoping a few hit it big. It’s not about avoiding failure; failure is baked into the model. The aim is to make sure the winners pay for all the losers—and then some. But in the Product Bubble, the stakes were raised. The flood of capital created immense pressure to deliver outsized returns. And delivering those returns required chasing unicorns. Unicorns are extraordinary by definition: billion-dollar companies born from ideas that often seemed ordinary at first glance. The problem was that creating them on such a large scale wasn’t sustainable. At its peak, VC firms were deploying $300 billion annually into startups. The math behind their portfolios was brutal: 10% of investments would be home runs, generating most of the returns.30% would break even or deliver modest returns.60% would fail outright. To meet these odds, VCs had to make a lot of bets. And when there’s too much money chasing too few good ideas, standards inevitably slip. Marginal ideas—the ones that wouldn’t have made the cut in leaner times—got funded. Suddenly, the market was full of startups promising billion-dollar visions, many of which had no business being there. To understand how this played out, consider the kinds of businesses VCs funded: A SaaS analytics company with $70 million in annual recurring revenue? That’s plausible.A government contractor with $80 million in EBITDA? Also plausible, with the right patience.A social network with 5 million daily users? Feasible with viral growth. But during the Product Bubble, it wasn’t enough to build a solid business. To attract VC attention, startups had to promise hyper-growth. That meant taking risks, stretching visions, and often pushing boundaries in ways that led to unintended consequences. The defining feature of this era was hyper-scale disruption. Startups weren’t just trying to compete; they were trying to upend entire industries. Uber and Airbnb are classic examples. Both started with seemingly innocent ideas—ride-sharing and room-sharing—but grew into juggernauts that reshaped cities, industries, and livelihoods. This kind of disruption had a dark side. Regulations were treated as obstacles to be skirted. Workers became disposable cogs. Quality and sustainability often took a back seat to growth. Consider the gig economy companies that flooded markets with subsidized pricing to undercut competitors. It worked—for a while. But when funding dried up, these models collapsed, leaving industries in disarray and consumers with fewer choices. The Product Bubble is filled with stories of founders whose good intentions led to unintended consequences: A food delivery service grew into a platform that strained restaurants and exploited drivers.A room-sharing idea morphed into Airbnb, accused of driving up rents and displacing residents.A social app for friends turned into Facebook, implicated in misinformation and polarization. These founders didn’t set out to harm anyone. They were chasing growth. But in doing so, they often overlooked the broader impacts of their decisions. By the end of the bubble, the unicorns of the era—Uber, Airbnb, and others—achieved valuations of $100 billion or more. But their success came at a cost: skyrocketing rents, increased congestion, and precarious working conditions. The relentless pursuit of scale had outstripped the ability to manage it responsibly. Why did this happen? Because the system was optimized for growth at all costs. Startups were incentivized to over-promise and under-deliver, cutting corners and sometimes ignoring regulations entirely. It worked, for a time. But now the bill is due. The Product Bubble has burst, thanks in part to rising interest rates and a cooling economy. This is a moment for everyone—founders, investors, and communities—to reflect. Does every idea need to become a billion-dollar company? Is chasing the biggest possible returns always the best approach? The answers seem increasingly clear: no. The future of venture capital doesn’t have to look like the Product Bubble. The focus can shift toward sustainable growth, responsible innovation, and long-term value. Founders can be encouraged to build resilient businesses that solve real problems, rather than chasing inflated valuations. The Product Bubble was a lesson in excess. It showed what happens when growth becomes an end in itself. But as the dust settles, there’s an opportunity to chart a better course. Future startups can be built on foundations of resilience, responsibility, and true value. The next era of innovation doesn’t have to be a bubble. It can be something better.]]></summary></entry><entry><title type="html">From Physics to Machine Learning: A Nobel Prize Worthy Journey</title><link href="https://asjad99.github.io/blog/2024/10/from-physics-to-machine-learning-a-nobel-prize-worthy-journey/" rel="alternate" type="text/html" title="From Physics to Machine Learning: A Nobel Prize Worthy Journey"/><published>2024-10-08T23:02:13+00:00</published><updated>2024-10-08T23:02:13+00:00</updated><id>https://asjad99.github.io/blog/2024/10/from-physics-to-machine-learning-a-nobel-prize-worthy-journey</id><content type="html" xml:base="https://asjad99.github.io/blog/2024/10/from-physics-to-machine-learning-a-nobel-prize-worthy-journey/"><![CDATA[<p>The 2024 Nobel Prize in Physics has been awarded to two pioneers whose groundbreaking work bridged physics and artificial intelligence, laying the foundation for modern artificial neural networks (ANNs). John J. Hopfield and Geoffrey E. Hinton didn’t just transform AI; they illuminated the profound and surprising connections between physics, biology, and machine learning. Their journey—spanning decades—offers a fascinating look at how ideas from one field can ignite revolutions in another.</p> <p>In 1982, physicist John J. Hopfield introduced a neural network model that mimicked the brain’s associative memory. His Hopfield Network could store patterns and recall them from incomplete inputs—like recognizing a familiar face from a blurry photograph. The brilliance of Hopfield’s approach lay in applying his expertise in statistical physics, specifically the behavior of spin glasses, a class of disordered magnetic materials.</p> <p>Hopfield’s network was more than a clever analogy; it was a bridge between physics and computation. He showed that the energy minimization principles governing physical systems could also describe how neural networks find stable states. This idea gave neural networks a formal mathematical foundation and revealed an elegant symmetry: both magnetic systems and neural networks seek to minimize energy, settling into configurations that make the most sense given their constraints.</p> <p>Geoffrey Hinton, one of the "godfathers" of deep learning, expanded on Hopfield’s ideas in the 1980s. Together with collaborators, Hinton developed the Boltzmann Machine, a probabilistic model that introduced randomness to neural networks. By drawing on the Boltzmann distribution from thermodynamics, Hinton created a system capable of tackling more complex learning problems.</p> <p>The Boltzmann Machine laid critical groundwork for deep learning. Hinton’s later development of the Restricted Boltzmann Machine (RBM) simplified these models and paved the way for today’s deep learning architectures. These breakthroughs made it possible to train deep, multilayered networks, unlocking the potential for AI to process vast amounts of data and uncover hidden patterns.</p> <p>The beauty of Hopfield and Hinton’s contributions lies in their interdisciplinary reach. Hopfield’s networks paralleled spin glass systems in physics, where particles settle into stable configurations by minimizing energy. The mathematical elegance extended even further: the Lyapunov function in Hopfield networks resembles the risk minimization strategies of portfolio theory in finance, highlighting the versatility of their ideas.</p> <p>Hinton’s deep learning innovations found applications across physics, biology, and beyond. In quantum mechanics, these models helped predict quantum phase transitions. In high-energy physics, they enabled particle detection in collider data. Hinton’s work also laid the groundwork for convolutional neural networks (CNNs), key to image recognition systems that now drive facial recognition, autonomous vehicles, and AI-powered medical diagnostics.</p> <figure class="kg-card kg-image-card kg-card-hascaption"><img src="/assets/img/ghost/9a1f1231e8-image.png" class="kg-image" alt="" loading="lazy" width="800" height="618" srcset="/assets/img/ghost/90501b032d-image.png 600w, /assets/img/ghost/9a1f1231e8-image.png 800w" sizes="(min-width: 720px) 720px"/><figcaption><b><strong style="white-space: pre-wrap;">Applications Across Physics, Biology, and Finance</strong></b></figcaption></figure> <p>The Nobel Committee’s recognition of Hopfield and Hinton underscores the transformative power of multidisciplinary thinking. Their work demonstrates how concepts from physics can revolutionize AI and, in turn, impact countless industries—from healthcare to transportation.</p> <p>This award serves as a powerful reminder: progress often comes from connecting ideas across fields. Hopfield and Hinton’s journey exemplifies how breakthroughs arise not just from technical mastery, but from a willingness to explore uncharted territory at the intersections of disciplines. By doing so, they paved the way for technologies like AlphaFold’s protein structure predictions and self-driving cars.</p> <p>As science and technology become ever more interconnected, the lessons of Hopfield and Hinton’s work resonate more than ever. Their achievements highlight the importance of thinking across boundaries, combining insights from multiple disciplines to tackle complex challenges.</p> <p>In celebrating their Nobel Prize, we’re not just honoring their past contributions. We’re recognizing the roadmap they’ve provided for future innovators: to see the patterns others miss, to connect the dots between disparate fields, and to push the boundaries of what’s possible.</p> <hr/> <p><em>This article was written in collaboration with LLM based writing assistants. </em></p>]]></content><author><name></name></author><summary type="html"><![CDATA[The 2024 Nobel Prize in Physics has been awarded to two pioneers whose groundbreaking work bridged physics and artificial intelligence, laying the foundation for modern artificial neural networks (ANNs). John J. Hopfield and Geoffrey E. Hinton didn’t just transform AI; they illuminated the profound and surprising connections between physics, biology, and machine learning. Their journey—spanning decades—offers a fascinating look at how ideas from one field can ignite revolutions in another. In 1982, physicist John J. Hopfield introduced a neural network model that mimicked the brain’s associative memory. His Hopfield Network could store patterns and recall them from incomplete inputs—like recognizing a familiar face from a blurry photograph. The brilliance of Hopfield’s approach lay in applying his expertise in statistical physics, specifically the behavior of spin glasses, a class of disordered magnetic materials. Hopfield’s network was more than a clever analogy; it was a bridge between physics and computation. He showed that the energy minimization principles governing physical systems could also describe how neural networks find stable states. This idea gave neural networks a formal mathematical foundation and revealed an elegant symmetry: both magnetic systems and neural networks seek to minimize energy, settling into configurations that make the most sense given their constraints. Geoffrey Hinton, one of the "godfathers" of deep learning, expanded on Hopfield’s ideas in the 1980s. Together with collaborators, Hinton developed the Boltzmann Machine, a probabilistic model that introduced randomness to neural networks. By drawing on the Boltzmann distribution from thermodynamics, Hinton created a system capable of tackling more complex learning problems. The Boltzmann Machine laid critical groundwork for deep learning. Hinton’s later development of the Restricted Boltzmann Machine (RBM) simplified these models and paved the way for today’s deep learning architectures. These breakthroughs made it possible to train deep, multilayered networks, unlocking the potential for AI to process vast amounts of data and uncover hidden patterns. The beauty of Hopfield and Hinton’s contributions lies in their interdisciplinary reach. Hopfield’s networks paralleled spin glass systems in physics, where particles settle into stable configurations by minimizing energy. The mathematical elegance extended even further: the Lyapunov function in Hopfield networks resembles the risk minimization strategies of portfolio theory in finance, highlighting the versatility of their ideas. Hinton’s deep learning innovations found applications across physics, biology, and beyond. In quantum mechanics, these models helped predict quantum phase transitions. In high-energy physics, they enabled particle detection in collider data. Hinton’s work also laid the groundwork for convolutional neural networks (CNNs), key to image recognition systems that now drive facial recognition, autonomous vehicles, and AI-powered medical diagnostics. Applications Across Physics, Biology, and Finance The Nobel Committee’s recognition of Hopfield and Hinton underscores the transformative power of multidisciplinary thinking. Their work demonstrates how concepts from physics can revolutionize AI and, in turn, impact countless industries—from healthcare to transportation. This award serves as a powerful reminder: progress often comes from connecting ideas across fields. Hopfield and Hinton’s journey exemplifies how breakthroughs arise not just from technical mastery, but from a willingness to explore uncharted territory at the intersections of disciplines. By doing so, they paved the way for technologies like AlphaFold’s protein structure predictions and self-driving cars. As science and technology become ever more interconnected, the lessons of Hopfield and Hinton’s work resonate more than ever. Their achievements highlight the importance of thinking across boundaries, combining insights from multiple disciplines to tackle complex challenges. In celebrating their Nobel Prize, we’re not just honoring their past contributions. We’re recognizing the roadmap they’ve provided for future innovators: to see the patterns others miss, to connect the dots between disparate fields, and to push the boundaries of what’s possible. This article was written in collaboration with LLM based writing assistants.]]></summary></entry><entry><title type="html">Responsible AI: Adopting Microsoft’s Framework for Ethical AI Use in Machine Learning Projects</title><link href="https://asjad99.github.io/blog/2024/10/untitled-4/" rel="alternate" type="text/html" title="Responsible AI: Adopting Microsoft’s Framework for Ethical AI Use in Machine Learning Projects"/><published>2024-10-08T09:13:24+00:00</published><updated>2024-10-08T09:13:24+00:00</updated><id>https://asjad99.github.io/blog/2024/10/untitled-4</id><content type="html" xml:base="https://asjad99.github.io/blog/2024/10/untitled-4/"><![CDATA[<h3></h3> <p>Artificial Intelligence (AI) holds immense potential to improve lives, drive efficiency, and offer solutions to some of society's biggest challenges. But with great power comes great responsibility. As AI systems are increasingly integrated into our daily lives, the importance of ensuring that these systems are developed and used responsibly cannot be overstated. Microsoft has taken a significant step in this direction with its "Responsible AI Standard v2," which provides a structured framework for the ethical use of AI. Let's dive into how adopting this framework can help promote responsible AI in machine learning projects, and how data scientists can leverage open-source tools to achieve this.</p> <h4 id="the-foundation-of-microsofts-responsible-ai-standard">The Foundation of Microsoft's Responsible AI Standard</h4> <p>Microsoft's Responsible AI Standard is the culmination of years of research, collaboration, and refinement, aimed at addressing the unique risks AI presents to society. The framework is based on six foundational goals that ensure accountability, transparency, fairness, reliability and safety, privacy, and inclusiveness. By operationalizing these principles, Microsoft aims to provide actionable guidance to developers, ensuring AI systems are ethical, inclusive, and safe for all users.</p> <h5 id="1-accountability-goals">1. <strong>Accountability Goals</strong></h5> <p>Accountability in AI involves taking ownership of the impact AI systems may have on individuals, organizations, and society. Microsoft's framework emphasizes the need for Impact Assessments during the design phase, documenting risks, defining acceptable uses, and ensuring stakeholders are informed. Human oversight is a core tenet, meaning there should always be responsible individuals overseeing the deployment and monitoring of AI systems to prevent adverse outcomes.</p> <p><strong>Concrete Adoption Example</strong>: In a machine learning project, data scientists can use MLOps practices to ensure accountability. Tools like <strong>MLflow</strong> or <strong>DVC</strong> (Data Version Control) can be used to register models, track experiments, and manage versions in development, testing, and production environments. Set up automated notifications for key events, such as model registration and data drift detection, to maintain full visibility and accountability across the model lifecycle. These practices help ensure that every change to a model is tracked and documented, promoting accountability.</p> <h5 id="2-transparency-goals">2. <strong>Transparency Goals</strong></h5> <p>Transparency in AI systems is critical to fostering trust. Microsoft's standard requires that stakeholders can understand how and why AI systems arrive at their decisions. This includes providing intelligibility for decision-making and clearly communicating the system's capabilities, limitations, and performance.</p> <p><strong>Concrete Adoption Example</strong>: Data scientists can use <strong>InterpretML</strong>, an open-source library, to develop AI explainer tools that provide human-understandable explanations for model predictions. This tool offers both global and local explanations of model behavior, enabling stakeholders to see which features affect the overall performance of a model or understand why an individual prediction was made. Incorporating these explanations into a model assessment workflow makes it easier for stakeholders to trust the AI system.</p> <h5 id="3-fairness-goals">3. <strong>Fairness Goals</strong></h5> <p>Fairness aims to reduce bias and prevent unequal treatment of demographic groups, especially marginalized communities. Microsoft emphasizes the evaluation of data sets for inclusiveness, continuous reassessment of system designs, and mitigating disparities in service quality and resource allocation.</p> <p><strong>Concrete Adoption Example</strong>: To ensure fairness in machine learning models, data scientists can use <strong>Fairlearn</strong>, an open-source toolkit that helps assess and improve fairness in AI systems. Fairlearn allows users to evaluate metrics across demographic subgroups, identify biases, and take steps to mitigate them. By analyzing the distribution of prediction values and performance metrics, data scientists can assess model effectiveness and ensure there is no significant disparity between groups, thereby providing fair and equitable outcomes.</p> <h5 id="4-reliability-safety-goals">4. <strong>Reliability &amp; Safety Goals</strong></h5> <p>Ensuring AI systems operate safely and consistently is a foundational aspect of Microsoft's Responsible AI Standard. This goal focuses on evaluating the conditions and settings where AI is deployed to ensure the system behaves as expected.</p> <p><strong>Concrete Adoption Example</strong>: Use <strong>Error Analysis</strong>, a tool that provides insights into how errors are distributed across a dataset. The <strong>Error Analysis</strong> toolkit can help data scientists identify problematic subgroups or data subsets with higher error rates than the overall benchmark, allowing targeted improvements to be made to increase the reliability of the AI system. Integrate this tool into a performance analysis dashboard to conduct error analysis and quickly identify areas for improvement.</p> <h5 id="5-privacy-security-goals">5. <strong>Privacy &amp; Security Goals</strong></h5> <p>Data privacy is at the heart of any responsible AI system. Microsoft's framework ensures that AI systems comply with privacy standards and handle user data securely. By embedding strong data governance practices, including assessing data quantity and quality, organizations can mitigate risks related to data misuse or breaches.</p> <p><strong>Concrete Adoption Example</strong>: Data scientists can leverage open-source tools like <strong>PySyft</strong> to work on privacy-preserving machine learning. Techniques such as differential privacy, federated learning, and secure multi-party computation help ensure that sensitive data remains secure during model training. Additionally, using encryption techniques and role-based access control for data ensures that data privacy is maintained at all stages of the project lifecycle.</p> <h5 id="6-inclusiveness-goals">6. <strong>Inclusiveness Goals</strong></h5> <p>AI should be designed with the goal of inclusiveness to benefit everyone. Microsoft emphasizes compliance with accessibility standards to ensure that AI systems can be used by as many people as possible, including those with disabilities.</p> <p><strong>Concrete Adoption Example</strong>: Data scientists can evaluate the inclusiveness of their models using the <strong>Responsible AI Dashboard</strong>, which integrates tools like <strong>Data Balance</strong> for understanding feature distributions and identifying any imbalances in data representation. Ensuring that features and outcomes are balanced across diverse user groups can help provide equitable results. Additionally, adherence to accessibility standards and involving affected communities in the development process can ensure that AI systems are accessible to individuals with disabilities.</p> <h4 id="leveraging-open-source-tools-for-responsible-ai">Leveraging Open Source Tools for Responsible AI</h4> <p>To support data scientists in adopting responsible AI practices, Microsoft provides several open-source tools that can be easily integrated into existing workflows:</p> <ul><li><strong>Fairlearn</strong>: For fairness assessment, Fairlearn helps identify groups that may be disproportionately negatively impacted by an AI system.</li><li><strong>InterpretML</strong>: To support transparency, InterpretML provides explanations for both global model behavior and individual predictions.</li><li><strong>Error Analysis</strong>: Enables data scientists to conduct detailed error analysis to identify high-error cohorts and improve model performance.</li><li><strong>DiCE</strong>: For counterfactual analysis, DiCE shows feature-perturbed versions of data points that could lead to different outcomes, providing insights into how changes in input can lead to desired results.</li><li><strong>EconML</strong>: For causal analysis, EconML helps answer “What If” questions to support data-driven decision-making.</li><li><strong>Responsible AI Dashboard</strong>: Integrates these tools, allowing data scientists to create comprehensive, customizable dashboards for end-to-end debugging, model assessment, and decision-making.</li></ul> <h4 id="applying-microsofts-framework-in-practice">Applying Microsoft's Framework in Practice</h4> <p>Adopting Microsoft’s Responsible AI framework involves integrating ethical considerations from the very beginning of AI development. Data scientists can start by conducting a thorough impact assessment before coding takes place, outlining potential risks to different demographic groups and ensuring that any sensitive uses of AI are thoroughly evaluated. By documenting these impact assessments, it becomes easier to update and address any issues that arise during subsequent releases.</p> <p>Continuous monitoring and evaluation are also crucial for AI systems. Using tools like <strong>Fairlearn</strong>, <strong>Error Analysis</strong>, and <strong>Responsible AI Dashboard</strong> ensures that AI systems evolve safely and responsibly as they interact with users and adapt to new contexts.</p> <h4 id="why-responsible-ai-matters">Why Responsible AI Matters</h4> <p>The adoption of responsible AI practices is essential to mitigating the risks of unintended consequences, discrimination, or harm. AI systems are making more decisions that affect people’s lives, such as determining loan eligibility, employment opportunities, and healthcare diagnostics. In such situations, the consequences of poorly designed or biased AI can be severe and far-reaching.</p> <p>By adopting Microsoft’s Responsible AI Standard and leveraging open-source tools, data scientists can ensure their AI systems are trustworthy, fair, and designed to minimize harm. This approach not only helps comply with regulations and standards but also builds public trust—a crucial element for widespread AI adoption.</p> <h4 id="final-thoughts">Final Thoughts</h4> <p>Microsoft's Responsible AI Standard offers a comprehensive framework to guide ethical AI development. Accountability, transparency, fairness, reliability, privacy, and inclusiveness are not just buzzwords but principles that must be integrated into every stage of AI design and deployment. By leveraging open-source tools like <strong>Fairlearn</strong>, <strong>InterpretML</strong>, <strong>DiCE</strong>, and the <strong>Responsible AI Dashboard</strong>, data scientists can create machine learning models that are not only effective but also ethical and equitable.</p> <p>Let’s build a future where AI not only solves problems but does so ethically, inclusively, and responsibly.</p> <p></p> <hr/> <p></p> <p>This post was written in collaboration with LLM based writing assist tools. </p>]]></content><author><name></name></author><summary type="html"><![CDATA[Artificial Intelligence (AI) holds immense potential to improve lives, drive efficiency, and offer solutions to some of society's biggest challenges. But with great power comes great responsibility. As AI systems are increasingly integrated into our daily lives, the importance of ensuring that these systems are developed and used responsibly cannot be overstated. Microsoft has taken a significant step in this direction with its "Responsible AI Standard v2," which provides a structured framework for the ethical use of AI. Let's dive into how adopting this framework can help promote responsible AI in machine learning projects, and how data scientists can leverage open-source tools to achieve this. The Foundation of Microsoft's Responsible AI Standard Microsoft's Responsible AI Standard is the culmination of years of research, collaboration, and refinement, aimed at addressing the unique risks AI presents to society. The framework is based on six foundational goals that ensure accountability, transparency, fairness, reliability and safety, privacy, and inclusiveness. By operationalizing these principles, Microsoft aims to provide actionable guidance to developers, ensuring AI systems are ethical, inclusive, and safe for all users. 1. Accountability Goals Accountability in AI involves taking ownership of the impact AI systems may have on individuals, organizations, and society. Microsoft's framework emphasizes the need for Impact Assessments during the design phase, documenting risks, defining acceptable uses, and ensuring stakeholders are informed. Human oversight is a core tenet, meaning there should always be responsible individuals overseeing the deployment and monitoring of AI systems to prevent adverse outcomes. Concrete Adoption Example: In a machine learning project, data scientists can use MLOps practices to ensure accountability. Tools like MLflow or DVC (Data Version Control) can be used to register models, track experiments, and manage versions in development, testing, and production environments. Set up automated notifications for key events, such as model registration and data drift detection, to maintain full visibility and accountability across the model lifecycle. These practices help ensure that every change to a model is tracked and documented, promoting accountability. 2. Transparency Goals Transparency in AI systems is critical to fostering trust. Microsoft's standard requires that stakeholders can understand how and why AI systems arrive at their decisions. This includes providing intelligibility for decision-making and clearly communicating the system's capabilities, limitations, and performance. Concrete Adoption Example: Data scientists can use InterpretML, an open-source library, to develop AI explainer tools that provide human-understandable explanations for model predictions. This tool offers both global and local explanations of model behavior, enabling stakeholders to see which features affect the overall performance of a model or understand why an individual prediction was made. Incorporating these explanations into a model assessment workflow makes it easier for stakeholders to trust the AI system. 3. Fairness Goals Fairness aims to reduce bias and prevent unequal treatment of demographic groups, especially marginalized communities. Microsoft emphasizes the evaluation of data sets for inclusiveness, continuous reassessment of system designs, and mitigating disparities in service quality and resource allocation. Concrete Adoption Example: To ensure fairness in machine learning models, data scientists can use Fairlearn, an open-source toolkit that helps assess and improve fairness in AI systems. Fairlearn allows users to evaluate metrics across demographic subgroups, identify biases, and take steps to mitigate them. By analyzing the distribution of prediction values and performance metrics, data scientists can assess model effectiveness and ensure there is no significant disparity between groups, thereby providing fair and equitable outcomes. 4. Reliability &amp; Safety Goals Ensuring AI systems operate safely and consistently is a foundational aspect of Microsoft's Responsible AI Standard. This goal focuses on evaluating the conditions and settings where AI is deployed to ensure the system behaves as expected. Concrete Adoption Example: Use Error Analysis, a tool that provides insights into how errors are distributed across a dataset. The Error Analysis toolkit can help data scientists identify problematic subgroups or data subsets with higher error rates than the overall benchmark, allowing targeted improvements to be made to increase the reliability of the AI system. Integrate this tool into a performance analysis dashboard to conduct error analysis and quickly identify areas for improvement. 5. Privacy &amp; Security Goals Data privacy is at the heart of any responsible AI system. Microsoft's framework ensures that AI systems comply with privacy standards and handle user data securely. By embedding strong data governance practices, including assessing data quantity and quality, organizations can mitigate risks related to data misuse or breaches. Concrete Adoption Example: Data scientists can leverage open-source tools like PySyft to work on privacy-preserving machine learning. Techniques such as differential privacy, federated learning, and secure multi-party computation help ensure that sensitive data remains secure during model training. Additionally, using encryption techniques and role-based access control for data ensures that data privacy is maintained at all stages of the project lifecycle. 6. Inclusiveness Goals AI should be designed with the goal of inclusiveness to benefit everyone. Microsoft emphasizes compliance with accessibility standards to ensure that AI systems can be used by as many people as possible, including those with disabilities. Concrete Adoption Example: Data scientists can evaluate the inclusiveness of their models using the Responsible AI Dashboard, which integrates tools like Data Balance for understanding feature distributions and identifying any imbalances in data representation. Ensuring that features and outcomes are balanced across diverse user groups can help provide equitable results. Additionally, adherence to accessibility standards and involving affected communities in the development process can ensure that AI systems are accessible to individuals with disabilities. Leveraging Open Source Tools for Responsible AI To support data scientists in adopting responsible AI practices, Microsoft provides several open-source tools that can be easily integrated into existing workflows: Fairlearn: For fairness assessment, Fairlearn helps identify groups that may be disproportionately negatively impacted by an AI system.InterpretML: To support transparency, InterpretML provides explanations for both global model behavior and individual predictions.Error Analysis: Enables data scientists to conduct detailed error analysis to identify high-error cohorts and improve model performance.DiCE: For counterfactual analysis, DiCE shows feature-perturbed versions of data points that could lead to different outcomes, providing insights into how changes in input can lead to desired results.EconML: For causal analysis, EconML helps answer “What If” questions to support data-driven decision-making.Responsible AI Dashboard: Integrates these tools, allowing data scientists to create comprehensive, customizable dashboards for end-to-end debugging, model assessment, and decision-making. Applying Microsoft's Framework in Practice Adopting Microsoft’s Responsible AI framework involves integrating ethical considerations from the very beginning of AI development. Data scientists can start by conducting a thorough impact assessment before coding takes place, outlining potential risks to different demographic groups and ensuring that any sensitive uses of AI are thoroughly evaluated. By documenting these impact assessments, it becomes easier to update and address any issues that arise during subsequent releases. Continuous monitoring and evaluation are also crucial for AI systems. Using tools like Fairlearn, Error Analysis, and Responsible AI Dashboard ensures that AI systems evolve safely and responsibly as they interact with users and adapt to new contexts. Why Responsible AI Matters The adoption of responsible AI practices is essential to mitigating the risks of unintended consequences, discrimination, or harm. AI systems are making more decisions that affect people’s lives, such as determining loan eligibility, employment opportunities, and healthcare diagnostics. In such situations, the consequences of poorly designed or biased AI can be severe and far-reaching. By adopting Microsoft’s Responsible AI Standard and leveraging open-source tools, data scientists can ensure their AI systems are trustworthy, fair, and designed to minimize harm. This approach not only helps comply with regulations and standards but also builds public trust—a crucial element for widespread AI adoption. Final Thoughts Microsoft's Responsible AI Standard offers a comprehensive framework to guide ethical AI development. Accountability, transparency, fairness, reliability, privacy, and inclusiveness are not just buzzwords but principles that must be integrated into every stage of AI design and deployment. By leveraging open-source tools like Fairlearn, InterpretML, DiCE, and the Responsible AI Dashboard, data scientists can create machine learning models that are not only effective but also ethical and equitable. Let’s build a future where AI not only solves problems but does so ethically, inclusively, and responsibly. This post was written in collaboration with LLM based writing assist tools.]]></summary></entry><entry><title type="html">Google Gemini updates: Flash 1.5, Gemma 2 and Project Astra</title><link href="https://asjad99.github.io/blog/2024/google-gemini-updates-flash-15-gemma-2-and-project-astra/" rel="alternate" type="text/html" title="Google Gemini updates: Flash 1.5, Gemma 2 and Project Astra"/><published>2024-05-14T00:00:00+00:00</published><updated>2024-05-14T00:00:00+00:00</updated><id>https://asjad99.github.io/blog/2024/google-gemini-updates-flash-15-gemma-2-and-project-astra</id><content type="html" xml:base="https://asjad99.github.io/blog/2024/google-gemini-updates-flash-15-gemma-2-and-project-astra/"><![CDATA[<p>Gemini breaks new ground with a faster model, longer context, AI agents and moreLearn more:Learn more:Learn more:Models &amp; ResearchProductsInfrastructure &amp; cloudTechnology Learn more: ProductsPlatformsDevices Learn more: Outreach &amp; initiativesLeadershipInside Google Learn more: May 14, 2024 We’re introducing a series of updates across the Gemini family of models, including the new 1.5 Flash, our lightweight model for speed and efficiency, and Project Astra, our vision for the future of AI assistants. Demis HassabisCEO of Google DeepMind, on behalf of the Gemini teamIn December, we launched our first natively multimodal model Gemini 1.0 in three sizes: Ultra, Pro and Nano. Just a few months later we released 1.5 Pro, with enhanced performance and a breakthrough long context window of 1 million tokens.Developers and enterprise customers have been putting 1.5 Pro to use in incredible ways and finding its long context window, multimodal reasoning capabilities and impressive overall performance incredibly useful.We know from user feedback that some applications need lower latency and a lower cost to serve. This inspired us to keep innovating, so today, we’re introducing Gemini 1.5 Flash: a model that’s lighter-weight than 1.5 Pro, and designed to be fast and efficient to serve at scale.Both 1.5 Pro and 1.5 Flash are available in public preview with a 1 million token context window in Google AI Studio and Vertex AI. And now, 1.5 Pro is also available with a 2 million token context window via waitlist to developers using the API and to Google Cloud customers.We’re also introducing updates across the Gemini family of models, announcing our next generation of open models, Gemma 2, and sharing progress on the future of AI assistants, with Project Astra.Context lengths of leading foundation models compared with Gemini 1.5’s 2 million token capability1.5 Flash is the newest addition to the Gemini model family and the fastest Gemini model served in the API. It’s optimized for high-volume, high-frequency tasks at scale, is more cost-efficient to serve and features our breakthrough long context window.While it’s a lighter weight model than 1.5 Pro, it’s highly capable of multimodal reasoning across vast amounts of information and delivers impressive quality for its size.The new Gemini 1.5 Flash model is optimized for speed and efficiency, is highly capable of multimodal reasoning and features our breakthrough long context window.1.5 Flash excels at summarization, chat applications, image and video captioning, data extraction from long documents and tables, and more. This is because it’s been trained by 1.5 Pro through a process called “distillation,” where the most essential knowledge and skills from a larger model are transferred to a smaller, more efficient model.Read more about 1.5 Flash in our updated Gemini 1.5 technical report, on the Gemini technology page, and learn about 1.5 Flash’s availability and pricing.Over the last few months, we’ve significantly improved 1.5 Pro, our best model for general performance across a wide range of tasks.Beyond extending its context window to 2 million tokens, we’ve enhanced its code generation, logical reasoning and planning, multi-turn conversation, and audio and image understanding through data and algorithmic advances. We see strong improvements on public and internal benchmarks for each of these tasks.1.5 Pro can now follow increasingly complex and nuanced instructions, including ones that specify product-level behavior involving role, format and style. We’ve improved control over the model’s responses for specific use cases, like crafting the persona and response style of a chat agent or automating workflows through multiple function calls. And we’ve enabled users to steer model behavior by setting system instructions.We added audio understanding in the Gemini API and Google AI Studio, so 1.5 Pro can now reason across image and audio for videos uploaded in Google AI Studio. And we’re now integrating 1.5 Pro into Google products, including Gemini Advanced and in Workspace apps.Read more about 1.5 Pro in our updated Gemini 1.5 technical report and on the Gemini technology page.Gemini Nano is expanding beyond text-only inputs to include images as well. Starting with Pixel, applications using Gemini Nano with Multimodality will be able to understand the world the way people do — not just through text, but also through sight, sound and spoken language.Read more about Gemini 1.0 Nano on Android.Today, we’re also sharing a series of updates to Gemma, our family of open models built from the same research and technology used to create the Gemini models.We’re announcing Gemma 2, our next generation of open models for responsible AI innovation. Gemma 2 has a new architecture designed for breakthrough performance and efficiency, and will be available in new sizes.The Gemma family is also expanding with PaliGemma, our first vision-language model inspired by PaLI-3. And we’ve upgraded our Responsible Generative AI Toolkit with LLM Comparator for evaluating the quality of model responses.Read more on the Developer blog.As part of Google DeepMind’s mission to build AI responsibly to benefit humanity, we’ve always wanted to develop universal AI agents that can be helpful in everyday life. That’s why today, we’re sharing our progress in building the future of AI assistants with Project Astra (advanced seeing and talking responsive agent).To be truly useful, an agent needs to understand and respond to the complex and dynamic world just like people do — and take in and remember what it sees and hears to understand context and take action. It also needs to be proactive, teachable and personal, so users can talk to it naturally and without lag or delay.While we’ve made incredible progress developing AI systems that can understand multimodal information, getting response time down to something conversational is a difficult engineering challenge. Over the past few years, we’ve been working to improve how our models perceive, reason and converse to make the pace and quality of interaction feel more natural.Building on Gemini, we’ve developed prototype agents that can process information faster by continuously encoding video frames, combining the video and speech input into a timeline of events, and caching this information for efficient recall.By leveraging our leading speech models, we also enhanced how they sound, giving the agents a wider range of intonations. These agents can better understand the context they’re being used in, and respond quickly, in conversation.With technology like this, it’s easy to envision a future where people could have an expert AI assistant by their side, through a phone or glasses. And some of these capabilities are coming to Google products, like the Gemini app and web experience, later this year.We’ve made incredible progress so far with our family of Gemini models, and we’re always striving to advance the state-of-the-art even further. By investing in a relentless production line of innovation, we’re able to explore new ideas at the frontier, while also unlocking the possibility of new and exciting Gemini use cases.Learn more about Gemini and its capabilities.Collection Sign up for our newsletters with product updates, event information, special offers, and more. Done. Just one step more. Check your inbox to confirm your subscription.</p> <div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>You can also subscribe with a different email address.
  
        Your information will be used in accordance with Google's privacy policy. You may opt out at any time.
</code></pre></div></div>]]></content><author><name></name></author><category term="external-posts"/><category term="google"/><summary type="html"><![CDATA[We’re sharing updates across our Gemini family of models and a glimpse of Project Astra, our vision for the future of AI assistants.]]></summary></entry><entry><title type="html">Future of Data Driven Process Optimisation</title><link href="https://asjad99.github.io/blog/2024/02/process-mining-open-challenges/" rel="alternate" type="text/html" title="Future of Data Driven Process Optimisation"/><published>2024-02-09T04:46:10+00:00</published><updated>2024-02-09T04:46:10+00:00</updated><id>https://asjad99.github.io/blog/2024/02/process-mining-open-challenges</id><content type="html" xml:base="https://asjad99.github.io/blog/2024/02/process-mining-open-challenges/"><![CDATA[<p>How can organizations analyze and improve their processes today in the context of emerging technologies like GenAI etc? No matter what technologies are brewing in the labs, the story of process optimization will always starts from event log. Event logs represent raw data of executed processes, capturing what happened, when, and in what sequence. Process mining techniques turn these logs into insights, revealing inefficiencies, bottlenecks, and opportunities for optimization. But here's the catch: the logs don’t always tell the full story.</p> <p>The process mining community has spent decades refining tools for discovering, checking, and enhancing processes based on event logs. Yet these methods struggle when the logs are messy—full of gaps, noise, and ambiguities—or when the processes themselves are unstructured and complex, like those in healthcare. What you get from the algorithms in these cases is often what’s called a "spaghetti model"—a tangled mess that’s as hard to interpret as it is to act on.</p> <p>The fundamental issue is this: traditional process mining methods focus narrowly on observed behavior. They extract insights from what’s explicitly recorded but fail to reason beyond it. They lack what researchers call <strong>common-sense reasoning</strong>, the ability to infer relationships or patterns that aren't directly logged but are nonetheless obvious to a human observer. This limitation creates a gap in understanding—one that leaves process analysts guessing, abstracting, or even discarding key insights.</p> <h3 id="the-limits-of-process-logs">The Limits of Process Logs</h3> <p>Imagine trying to analyze a hospital’s clinical workflows. The event logs might tell you when a patient arrived, what tests were performed, and when they were discharged. But they probably won’t capture why certain decisions were made, the cascading effects of those decisions, or the informal handoffs and workarounds that happen in practice. Without this context, even the most sophisticated process mining algorithms can only provide a partial picture.</p> <p>And it’s not just about missing context. Logs are often incomplete or biased, skewed by errors in data recording or by systemic gaps in what’s logged in the first place. The result is a process model that might fit the data but misses the mark in reality. It’s like trying to reverse-engineer a symphony from a handful of notes—technically accurate, perhaps, but far from complete.</p> <p>This limitation becomes even more pronounced in complex domains like healthcare, where understanding process behavior isn’t just about efficiency—it’s about improving outcomes that directly affect people’s lives. The stakes are too high for guesswork or overly simplistic models.</p> <h3 id="toward-knowledge-centric-process-mining">Toward Knowledge-Centric Process Mining</h3> <p>If event logs alone can’t give us the full picture, what can? One promising answer lies in <strong>knowledge graphs</strong>. These structures represent relationships between entities—like people, processes, or decisions—in a way that’s both flexible and richly detailed. They’re not just about storing data; they’re about modeling knowledge, including the implicit relationships that traditional process mining methods overlook.</p> <p>Think of a knowledge graph as a map, where nodes represent concepts (e.g., “Patient,” “Test,” “Diagnosis”) and edges represent relationships between them (e.g., “undergoes,” “leads to,” “associated with”). This map can encode the hierarchical, cascading, and often abstract relationships that make real-world processes so complex. More importantly, it can provide the reasoning capabilities that traditional methods lack.</p> <p>For instance, a knowledge graph could help infer missing links in a process: “If a patient undergoes a test and the test indicates a risk, then a follow-up procedure is likely.” These kinds of inferences aren’t explicitly logged but are critical for understanding the broader dynamics of a process.</p> <h3 id="adding-intelligence-llms-and-knowledge-graphs">Adding Intelligence: LLMs and Knowledge Graphs</h3> <p>The other exciting development is the rise of <strong>Large Language Models (LLMs)</strong> fine-tuned on domain-specific data. Unlike traditional process mining algorithms, LLMs can process and interpret unstructured data—like text descriptions, clinical notes, or business rules—and integrate it into a structured understanding of the process. Paired with knowledge graphs, they can enrich process analytics with reasoning capabilities that go beyond pattern recognition.</p> <p>Imagine this scenario: You’re analyzing a hospital’s patient journey. The event logs give you a timeline of actions, but the LLM, trained on the organization’s documentation, adds context. It identifies that a delay in a specific test is linked to resource constraints and that these delays disproportionately affect certain patient groups. The knowledge graph then ties it all together, showing how these delays cascade through the system and identifying where interventions would have the most impact.</p> <p>This isn’t just about automating insights; it’s about augmenting human understanding. By combining the data-driven rigor of process mining with the contextual intelligence of LLMs and knowledge graphs, we can give process analysts tools that are both more powerful and more intuitive.</p> <h3 id="challenges-and-opportunities">Challenges and Opportunities</h3> <p>Of course, this vision comes with challenges. Building and maintaining knowledge graphs requires significant effort, especially in dynamic environments where processes and relationships are constantly evolving. And while LLMs are incredibly versatile, they’re not infallible—they’re only as good as the data they’re trained on, and they can introduce their own biases if not carefully managed.</p> <p>But the potential payoff is enormous. By bridging the gap between data and knowledge, we can move from simply observing processes to truly understanding them. This shift could transform how organizations approach everything from operational efficiency to customer experience, unlocking insights that were previously out of reach.</p> <h3 id="rethinking-process-analytics">Rethinking Process Analytics</h3> <p>Process analytics has always been about making sense of complexity. But as processes grow more intricate and the data we collect becomes more fragmented, the limitations of traditional methods become harder to ignore. The next frontier lies in embracing a <strong>knowledge-centric approach</strong>, one that integrates the strengths of process mining, knowledge graphs, and LLMs.</p> <p>This isn’t just an incremental improvement; it’s a rethinking of how we approach process discovery and optimization. It’s about recognizing that processes are more than the sum of their logged events—they’re dynamic, contextual, and deeply human. And understanding them fully requires tools that are just as dynamic, just as contextual, and just as human-centered.</p> <p>The future of process analytics isn’t about replacing the old tools; it’s about expanding what’s possible. By leveraging knowledge graphs and LLMs, we can move closer to that ideal—where process mining doesn’t just model what we’ve done but helps us imagine what we could do next.</p>]]></content><author><name></name></author><category term="#blog"/><summary type="html"><![CDATA[How can organizations analyze and improve their processes today in the context of emerging technologies like GenAI etc? No matter what technologies are brewing in the labs, the story of process optimization will always starts from event log. Event logs represent raw data of executed processes, capturing what happened, when, and in what sequence. Process mining techniques turn these logs into insights, revealing inefficiencies, bottlenecks, and opportunities for optimization. But here's the catch: the logs don’t always tell the full story. The process mining community has spent decades refining tools for discovering, checking, and enhancing processes based on event logs. Yet these methods struggle when the logs are messy—full of gaps, noise, and ambiguities—or when the processes themselves are unstructured and complex, like those in healthcare. What you get from the algorithms in these cases is often what’s called a "spaghetti model"—a tangled mess that’s as hard to interpret as it is to act on. The fundamental issue is this: traditional process mining methods focus narrowly on observed behavior. They extract insights from what’s explicitly recorded but fail to reason beyond it. They lack what researchers call common-sense reasoning, the ability to infer relationships or patterns that aren't directly logged but are nonetheless obvious to a human observer. This limitation creates a gap in understanding—one that leaves process analysts guessing, abstracting, or even discarding key insights. The Limits of Process Logs Imagine trying to analyze a hospital’s clinical workflows. The event logs might tell you when a patient arrived, what tests were performed, and when they were discharged. But they probably won’t capture why certain decisions were made, the cascading effects of those decisions, or the informal handoffs and workarounds that happen in practice. Without this context, even the most sophisticated process mining algorithms can only provide a partial picture. And it’s not just about missing context. Logs are often incomplete or biased, skewed by errors in data recording or by systemic gaps in what’s logged in the first place. The result is a process model that might fit the data but misses the mark in reality. It’s like trying to reverse-engineer a symphony from a handful of notes—technically accurate, perhaps, but far from complete. This limitation becomes even more pronounced in complex domains like healthcare, where understanding process behavior isn’t just about efficiency—it’s about improving outcomes that directly affect people’s lives. The stakes are too high for guesswork or overly simplistic models. Toward Knowledge-Centric Process Mining If event logs alone can’t give us the full picture, what can? One promising answer lies in knowledge graphs. These structures represent relationships between entities—like people, processes, or decisions—in a way that’s both flexible and richly detailed. They’re not just about storing data; they’re about modeling knowledge, including the implicit relationships that traditional process mining methods overlook. Think of a knowledge graph as a map, where nodes represent concepts (e.g., “Patient,” “Test,” “Diagnosis”) and edges represent relationships between them (e.g., “undergoes,” “leads to,” “associated with”). This map can encode the hierarchical, cascading, and often abstract relationships that make real-world processes so complex. More importantly, it can provide the reasoning capabilities that traditional methods lack. For instance, a knowledge graph could help infer missing links in a process: “If a patient undergoes a test and the test indicates a risk, then a follow-up procedure is likely.” These kinds of inferences aren’t explicitly logged but are critical for understanding the broader dynamics of a process. Adding Intelligence: LLMs and Knowledge Graphs The other exciting development is the rise of Large Language Models (LLMs) fine-tuned on domain-specific data. Unlike traditional process mining algorithms, LLMs can process and interpret unstructured data—like text descriptions, clinical notes, or business rules—and integrate it into a structured understanding of the process. Paired with knowledge graphs, they can enrich process analytics with reasoning capabilities that go beyond pattern recognition. Imagine this scenario: You’re analyzing a hospital’s patient journey. The event logs give you a timeline of actions, but the LLM, trained on the organization’s documentation, adds context. It identifies that a delay in a specific test is linked to resource constraints and that these delays disproportionately affect certain patient groups. The knowledge graph then ties it all together, showing how these delays cascade through the system and identifying where interventions would have the most impact. This isn’t just about automating insights; it’s about augmenting human understanding. By combining the data-driven rigor of process mining with the contextual intelligence of LLMs and knowledge graphs, we can give process analysts tools that are both more powerful and more intuitive. Challenges and Opportunities Of course, this vision comes with challenges. Building and maintaining knowledge graphs requires significant effort, especially in dynamic environments where processes and relationships are constantly evolving. And while LLMs are incredibly versatile, they’re not infallible—they’re only as good as the data they’re trained on, and they can introduce their own biases if not carefully managed. But the potential payoff is enormous. By bridging the gap between data and knowledge, we can move from simply observing processes to truly understanding them. This shift could transform how organizations approach everything from operational efficiency to customer experience, unlocking insights that were previously out of reach. Rethinking Process Analytics Process analytics has always been about making sense of complexity. But as processes grow more intricate and the data we collect becomes more fragmented, the limitations of traditional methods become harder to ignore. The next frontier lies in embracing a knowledge-centric approach, one that integrates the strengths of process mining, knowledge graphs, and LLMs. This isn’t just an incremental improvement; it’s a rethinking of how we approach process discovery and optimization. It’s about recognizing that processes are more than the sum of their logged events—they’re dynamic, contextual, and deeply human. And understanding them fully requires tools that are just as dynamic, just as contextual, and just as human-centered. The future of process analytics isn’t about replacing the old tools; it’s about expanding what’s possible. By leveraging knowledge graphs and LLMs, we can move closer to that ideal—where process mining doesn’t just model what we’ve done but helps us imagine what we could do next.]]></summary></entry><entry><title type="html">On Curse of Dimensionality</title><link href="https://asjad99.github.io/blog/2024/01/curse-of-dimensionality/" rel="alternate" type="text/html" title="On Curse of Dimensionality"/><published>2024-01-31T08:26:55+00:00</published><updated>2024-01-31T08:26:55+00:00</updated><id>https://asjad99.github.io/blog/2024/01/curse-of-dimensionality</id><content type="html" xml:base="https://asjad99.github.io/blog/2024/01/curse-of-dimensionality/"><![CDATA[<p>In this post, I will be sharing my understanding of feature selection process:</p> <p>In Theory feature selection can help in reducing the dimensionality, leading to simpler models that are easier to understand and interpret. Models with large number of features are also more prone to overfitting, especially if the dataset is not large enough to support the complexity. Simplifying the feature set can help in building more generalizable models. Overfitting is more likely to happen with a very large number of features, especially if some of those features do not contribute to the predictive power of the model. Models with fewer features are generally faster to train and require less computational resources.</p> <p><strong>But what about in Practice?</strong></p> <blockquote>The curse of dimensionality isn’t meaningful in practice because out space isn’t just a bunch of meaningless cartesian coordinates. We create structure, using trees, neural nets, etc. We regularize using bagging, weight decay, dropout, etc. We find that therefore we actually can add lots of columns without seeing problems in practice. - Jeremy Howard </blockquote> <p>More complex models with a large number of features can capture more intricate/complex patterns in the data. In practice, modern machine learning techniques have developed various methods to mitigate the issues arising from high-dimensional spaces. Modern algorithms don't treat the feature space as a simple Cartesian space but instead learn non-linear and complex boundaries within it. This structuring allows them to handle high-dimensional data more effectively than traditional statistical models. Ensemble methods like bagging can help prevent overfitting effectively and reduce the model's complexity, making it less sensitive to the curse of dimensionality. Although, assumption here is sufficiently large and diverse dataset to train on.</p> <p>While theoretically, a large number of features might pose challenges, in practice, if a model with many features consistently performs better on a well-constructed validation set, it's a strong argument in favor of using more features. Domain knowledge can also guide which features are likely to be relevant, which might not be immediately apparent through algorithmic feature selection techniques alone.</p> <p><strong>Performing Experiments for Feature Selection</strong>:</p> <p>This is beneficial when you suspect that not all features contribute equally to the predictive power of the model, or when you need to build a simpler, more interpretable mode. In practice, a common approach is to start with a model using all available features and then iteratively refine the feature set based on model performance and domain knowledge. This iterative approach helps in understanding the contribution of different features to the model and in finding a balance between model complexity and performance.</p> <p>In summary, while theoretical considerations about feature selection and the curse of dimensionality are important, practical machine learning often involves empirical testing and the application of advanced techniques to mitigate these issues. </p>]]></content><author><name></name></author><category term="Machine Learning"/><category term="#docs"/><summary type="html"><![CDATA[In this post, I will be sharing my understanding of feature selection process: In Theory feature selection can help in reducing the dimensionality, leading to simpler models that are easier to understand and interpret. Models with large number of features are also more prone to overfitting, especially if the dataset is not large enough to support the complexity. Simplifying the feature set can help in building more generalizable models. Overfitting is more likely to happen with a very large number of features, especially if some of those features do not contribute to the predictive power of the model. Models with fewer features are generally faster to train and require less computational resources. But what about in Practice? The curse of dimensionality isn’t meaningful in practice because out space isn’t just a bunch of meaningless cartesian coordinates. We create structure, using trees, neural nets, etc. We regularize using bagging, weight decay, dropout, etc. We find that therefore we actually can add lots of columns without seeing problems in practice. - Jeremy Howard More complex models with a large number of features can capture more intricate/complex patterns in the data. In practice, modern machine learning techniques have developed various methods to mitigate the issues arising from high-dimensional spaces. Modern algorithms don't treat the feature space as a simple Cartesian space but instead learn non-linear and complex boundaries within it. This structuring allows them to handle high-dimensional data more effectively than traditional statistical models. Ensemble methods like bagging can help prevent overfitting effectively and reduce the model's complexity, making it less sensitive to the curse of dimensionality. Although, assumption here is sufficiently large and diverse dataset to train on. While theoretically, a large number of features might pose challenges, in practice, if a model with many features consistently performs better on a well-constructed validation set, it's a strong argument in favor of using more features. Domain knowledge can also guide which features are likely to be relevant, which might not be immediately apparent through algorithmic feature selection techniques alone. Performing Experiments for Feature Selection: This is beneficial when you suspect that not all features contribute equally to the predictive power of the model, or when you need to build a simpler, more interpretable mode. In practice, a common approach is to start with a model using all available features and then iteratively refine the feature set based on model performance and domain knowledge. This iterative approach helps in understanding the contribution of different features to the model and in finding a balance between model complexity and performance. In summary, while theoretical considerations about feature selection and the curse of dimensionality are important, practical machine learning often involves empirical testing and the application of advanced techniques to mitigate these issues.]]></summary></entry></feed>