<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[Genetic Algorithms in R]]></title><description><![CDATA[A search heuristic and optimization technique inspired by Darwin's theory of natural selection. It is used to find optimal or near-optimal solutions to complex problems by evolving a population of candidate solutions over multiple generations.]]></description><link>https://sidiculouus-ga.hashnode.dev</link><image><url>https://cdn.hashnode.com/uploads/logos/69f2def0b18c97823374784c/21c2dabd-8722-4473-bd86-4d50a574e2fb.png</url><title>Genetic Algorithms in R</title><link>https://sidiculouus-ga.hashnode.dev</link></image><generator>RSS for Node</generator><lastBuildDate>Sun, 11 Oct 2026 08:29:23 GMT</lastBuildDate><atom:link href="https://sidiculouus-ga.hashnode.dev/rss.xml" rel="self" type="application/rss+xml"/><language><![CDATA[en]]></language><ttl>60</ttl><item><title><![CDATA[I Taught a Computer to Be a Lazy Traveller — Using Genetic Algorithms in R]]></title><description><![CDATA[There's a classic puzzle in computer science that sounds deceptively simple: given a list of cities, what's the shortest route that visits all of them exactly once and returns to the start?
That's the]]></description><link>https://sidiculouus-ga.hashnode.dev/i-taught-a-computer-to-be-a-lazy-traveller-using-genetic-algorithms-in-r</link><guid isPermaLink="true">https://sidiculouus-ga.hashnode.dev/i-taught-a-computer-to-be-a-lazy-traveller-using-genetic-algorithms-in-r</guid><category><![CDATA[genetic-algorithms]]></category><category><![CDATA[Machine Learning]]></category><category><![CDATA[R Programming]]></category><category><![CDATA[optimization]]></category><category><![CDATA[Data Science]]></category><dc:creator><![CDATA[Sidharth Vijayan]]></dc:creator><pubDate>Sat, 16 May 2026 19:57:42 GMT</pubDate><enclosure url="https://cdn.hashnode.com/uploads/covers/69f2def0b18c97823374784c/7093f30a-a570-493d-be24-535197dc4342.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>There's a classic puzzle in computer science that sounds deceptively simple: <em>given a list of cities, what's the shortest route that visits all of them exactly once and returns to the start?</em></p>
<p>That's the <strong>Travelling Salesman Problem</strong> (TSP). And the brutal truth? Nobody has found a perfect, fast solution for large inputs. It belongs to a class of problems that gets exponentially harder as you add more cities. For 10 cities there are over 3 million possible routes. For 30? Don't even try to count.</p>
<p>So for my add-on project in the Data Science with R Programming course, I decided to tackle it the way nature does — imperfectly, iteratively, but surprisingly well.</p>
<hr />
<h2>Why Genetic Algorithms?</h2>
<p>The honest answer: I was looking for something that <em>felt</em> smart without requiring me to brute-force my way through millions of possibilities.</p>
<p>Genetic Algorithms (GAs) are a family of optimization techniques inspired by biological evolution. The basic idea is that you start with a population of random solutions, evaluate how good each one is, and keep breeding the better ones together — occasionally throwing in a mutation — until the population collectively gets smarter. It's survival of the fittest, applied to route planning.</p>
<p>For TSP specifically, GAs made a lot of sense. There's no closed-form equation to optimize. The search space is combinatorial and chaotic. A GA doesn't need to know <em>why</em> a route is good — it just needs to measure <em>how good</em> it is, and work from there.</p>
<hr />
<h2>Doing This in R — A New Language for Me</h2>
<p>I'd been using Python for most of my data science work, so picking up R for this project was a deliberate challenge. R has a reputation for statistics and data analysis, not algorithm implementation. But working through this in R taught me that it's actually quite capable for custom algorithmic work — you just think about it differently.</p>
<p>The vector operations are intuitive once you stop fighting them. <code>sample()</code>, <code>which()</code>, <code>seq_along()</code> — these become second nature quickly. Writing the mutation step (randomly swapping two cities in a route) felt clean and readable in R in a way I didn't expect.</p>
<p>I also ran everything on <strong>Google Colab</strong>, which handled the R kernel without any setup headaches.</p>
<hr />
<h2>How the Algorithm Works</h2>
<p>Here's a breakdown of what I actually implemented, step by step:</p>
<h3>1. Initialization</h3>
<p>A population of random routes is generated. Each route is just a permutation of all the city indices — a possible order to visit them.</p>
<h3>2. Distance Calculation</h3>
<p>For each route, I calculate the total Euclidean distance between consecutive cities. This is the core metric everything else depends on.</p>
<h3>3. Fitness Evaluation</h3>
<p>Fitness is defined as:</p>
<pre><code class="language-plaintext">Fitness = 1 / (1 + Distance)
</code></pre>
<p>Shorter routes get higher fitness. Simple, but effective.</p>
<h3>4. Tournament Selection</h3>
<p>To pick parents for the next generation, I randomly pull three individuals from the population and take the best one. Run this twice, and you have two parents. This method avoids always picking the same top performers, which keeps diversity in the gene pool.</p>
<h3>5. Crossover</h3>
<p>A contiguous segment from Parent 1 is copied into the child. The remaining cities are filled in from Parent 2, in order, skipping any that are already present. This preserves valid permutations — no city gets visited twice.</p>
<h3>6. Mutation</h3>
<p>With a small probability, two cities in a route get swapped. This is what keeps the algorithm from getting stuck in a local optimum. A tiny nudge that occasionally opens up a better path.</p>
<h3>7. Elitism</h3>
<p>The single best solution from the current generation is always carried forward, unchanged. This ensures the algorithm never <em>loses</em> its best answer while exploring new ones.</p>
<p>Repeat steps 3–7 for hundreds of generations, and the population gradually converges on a near-optimal route.</p>
<hr />
<h2>Results</h2>
<p>I tested with two problem sizes:</p>
<p><strong>10 cities:</strong> The algorithm converged fast. Best distance landed around <strong>325.93</strong>. You could watch the fitness jump in the early generations and then plateau — the classic convergence curve.</p>
<p><strong>30 cities:</strong> Harder, slower, but still effective. Best distance reached approximately <strong>565.77</strong>. Convergence took longer, and the average population distance kept improving even when the best route had stabilised.</p>
<p>A few things I noticed in the results that I found genuinely interesting:</p>
<ul>
<li><p>The early generations improve <em>fast</em>. The random initial population is so bad that almost any crossover makes things better.</p>
</li>
<li><p>Mutation occasionally causes a sharp, sudden drop in best distance — a lucky swap that opens up a better sub-path.</p>
</li>
<li><p>Larger city counts don't just slow things down linearly. You really feel the combinatorial explosion.</p>
</li>
</ul>
<img src="https://cdn.hashnode.com/uploads/covers/69f2def0b18c97823374784c/c6cf070f-5d17-4eb7-b20d-b37d0ed6b826.png" alt="" style="display:block;margin:0 auto" />

<hr />
<h2>What I Actually Learned</h2>
<p>Beyond the algorithm itself, a few things stuck with me from this project:</p>
<p><strong>Fitness functions matter more than you'd think.</strong> Choosing <code>1 / (1 + distance)</code> instead of just <code>1 / distance</code> avoids division-by-zero weirdness and keeps the values in a range that's easy to work with. A small design decision, but it matters.</p>
<p><strong>Elitism is underrated.</strong> I initially ran the algorithm without it and watched good solutions disappear between generations. Adding elitism — just carrying forward the single best individual — made convergence visibly more stable.</p>
<p><strong>R is more versatile than its reputation suggests.</strong> I went in expecting to miss Python's libraries. I came out respecting R's expressiveness for this kind of work. It's not the language I'd reach for first for systems-level code, but for algorithmic experimentation with data? It holds up.</p>
<p><strong>Near-optimal is often good enough.</strong> This is maybe the most practically useful idea from the whole project. TSP is unsolvable "perfectly" at scale, but a solution that's 5% off optimal — found in seconds — is infinitely more useful than an exact answer that takes years to compute. GAs embody that trade-off well.</p>
<hr />
<h2>Where This Stuff Actually Gets Used</h2>
<p>TSP isn't just an academic puzzle. The same underlying problem — visit a set of locations in the best order — shows up constantly:</p>
<ul>
<li><p>Delivery route planning (think last-mile logistics)</p>
</li>
<li><p>Manufacturing: optimising the order in which a robot arm visits drill points</p>
</li>
<li><p>DNA sequencing: some assembly problems are structurally similar to TSP</p>
</li>
<li><p>Network packet routing</p>
</li>
<li><p>Robotics path planning</p>
</li>
</ul>
<p>Genetic Algorithms are used across all of these, usually as one layer in a larger optimization stack.</p>
<hr />
<h2>Final Thoughts</h2>
<p>I came into this project thinking "genetic algorithm" sounded more complicated than it was. In practice, the core ideas are accessible: start random, select the good stuff, mix it together, occasionally shake things up. It's the implementation details — crossover strategies, mutation rates, population sizing — where it gets interesting.</p>
<p>Running this in R pushed me to think carefully about the code rather than reach for a pre-built library. That friction was useful. I understood what I built.</p>
<p>If you want to poke around the actual code, it's here: <a href="https://colab.research.google.com/drive/1M5bwDfvXnBH3uHBl27xmw810X9Q2ubFA#scrollTo=x_ZPCWst5uYo"><strong>Google Colab Notebook →</strong></a></p>
<hr />
<p><em>Built as part of a Data Science with R Programming course project. First time writing a genetic algorithm from scratch.</em></p>
]]></content:encoded></item></channel></rss>