Thursday, May 6, 2021

Genetic algorithms vs. Genetic Programming

In the Introduction to Genetic Algorithms video tutorial, we saw how GAs are powerful evolutionary algorithms that are used to synthesize solutions and optimize parameters for problems that are too complex or convoluted for humans to solve on their own. 


They work well in cases where semantics isn’t a huge concern. So in the case of the video, each possible step the mouse was able to take was encoded as an integer (gene). Strings of these genes (chromosomes) were built, and we kept evolving them through mutation and crossover until we found a chromosome that got the mouse from the start of the maze over to the cheese at the end. 


We weren’t concerned about what order the integers were in a chromosome, or if any of the states were “invalid”. Sure, there were some chromosomes that had a lot of noise or useless actions, but nothing that inherently didn’t make sense. This is what I mean by “semantics”.

So what happens if we wanted to use GA to evolve a mathematical expression? An area where semantics does matter.


The goal here will be to find a Mathematical expression that will yield the following outputs from inputs:


When input = 2, Output = 9

When Input = 0, Output = 5

When Input = 1, Output = 6

When Input = .5, Output = 5.125


The exact expression we’re looking for the computer to find is y = x^2 + 5.


We could encode each mathematical operation as an integer:


0: x

1: ^

2: *

3: +

4: -

5 through 9: (can be a number/digit in the range 0 to 5)


A few possible states we could generate:


701846 (would be 2x^3 - 1)

601846 (would be x^3 - 1)

031111 (would be x+^^^^)

222223 (would be *****+)

434343 (would be -+-+-+)


The first two states aren’t exactly the winning expression that we want, but they’re valid. The last 3 though are completely bogus. 


Even if we started with a proper string, a slight alteration during the mutation process can have a drastic result on the offspring. A generated expression such as 701846 undergoing a mutation could result in 101846, which now results in the expression ^x^3 - 1 — which doesn’t make much sense, does it? 


In order to deal with these invalid states, we'd need to include mechanisms into the parser that would repair or the damaged expressions prior to evaluating their fitness. It's certainly doable, but not very practical, and it’s definitely not scalable as we advance to building more complicated expressions. In theory, we’d need to manually code in more and more rules on how the computer should repair/handle invalid states. 



Enter Genetic Programming



Genetic programming has qualities that naturally safeguards it from spitting out faulty states like above.


Unlike GA, genetic programs are typically represented as a structure of “nodes” connected to each other somehow in memory. One of the more common structures used to represent them are trees.


The math function (4 + (X/25) - (9 * sqrt(X)) represented by a tree structure:


A math function represented by a tree structure



Since we’re dealing with a tree data structure now, we won’t be able to just mutate the genes by flipping a bit, or swapping the value of an integer - more effort will need to go in to the mutation and crossover operations (if you're looking for specifics on the code part, I'll be covering this part in detail in future articles).


But the nice thing here is that the restrictions are built directly into the tree, so we don’t need to deal with invalid states.


The root node, and also the internal nodes (nodes with child nodes) can only be operators. The terminal nodes (those without child nodes) can only contain constants or variables. We don’t connect an operator to a constant or variable node. Connecting nodes in this manner prevents us from generating a state like +X/-9+6.


It’s possible we could run into divide by zero errors, but these can be easily dealt with during the fitness evaluation stage of the GP. 


The bottom line


Practically speaking, there’s a fine line between what constitutes as a “genetic algorithm” or “genetic program”. At the end of the day — it really comes down to the classification and context of the problem you’re solving, the particular mechanics you’re using, and the format of the states that you’re generating.


In general, the primary differences come down to this:


Genetic Algorithms are traditionally built using strings of bits, integers, or characters that represent the genes that make up a chromosome. While this can make it pretty easy to manage and and evolve a GA, it can make evaluating them more cumbersome when they are applied to applications where they have a propensity to generate invalid states. Additionally, the fact that GA chromosomes are typically restricted by a fixed length results in decreased flexibility in the ability to scale as a problem becomes more complex.


With Genetic Programs, we adopt the same underlying evolutionary concepts and foundation from GA, but encode the individuals with a tree (or similar) structure rather than a string or list of genes. A consequence of this is that the overhead required to grow and maintain these structures is significantly increased. On the positive side, we have more control to enforce limits directly into the programs we build which greatly decreases the occurrence of invalid states. Due to the variable size of this type of representation, GP individuals are inherently more flexible than GA chromosomes, and can grow and scale as needed according to a problem's complexity — while efficiently preserving limits and rules throughout the tree.

Thursday, February 18, 2021

Checking if a string is a permutation of a Palindrome


During this post we’ll be writing an algorithm that determines whether a string is a permutation of a Palindrome. We’re not talking about finding out whether a string is a Palindrome, but whether if a string is a permutation of a Palindrome. 


In other words, given a string of scrambled characters, we want to see if it would be possible to rearrange those characters into a word that spells the forward as it does backward (i.e: “oonn” can be re-arranged to form the palindrome “noon”). 


You might be tempted to write code that finds every possible permutation of characters to see if one forms a Palindrome. This would work for small strings but would quickly become time consuming as the amount of possible permutations rises dramatically. And in fact, it’s also not needed. This is likely one such reason why this question (and variations of it) emerge frequently in interview questions and other coding brainteasers — to test the ability for you to take a step back — assess a problem and carefully break it down into its most critical components to produce an efficient and optimized solution.


“Wow”, “Racecar”, “Radar”, “Level”, “Noon” are all examples of palindromes. They each have the following qualities in common: all have an even count of characters, or at most 1 character with an odd # of occurrences. This means that instead of going crazy generating every permutation of a string and checking whether each is a palindrome, we just need to check if the string we start with abides by these rules. 

 

Let’s apply these rules to the string “aacecrr” and see if we can accurately conclude whether the letters in “aacecrr” can be rearranged into a palindrome. We’ll keep count of the # of times we’ve seen each character, along with a count that keeps track of the odd # of characters — incrementing by 1 when any current character count is odd, and decrementing it by 1 when any current character count is even. Below is an illustrated breakdown of what we’ll do as we iterate through each letter in the string (the current character and the counts being decremented/incremented are highlighted in blue):


A A C E C R R


A: 1  (1 is odd so numOdd should be incremented by 1)

numOdd: 1


A A C E C R R


A: 2  (2 is even so numOdd should be decremented by 1)

numOdd: 0


A A C E C R R


A: 2

C: 1 (odd so increment numOdd by 1)

numOdd: 1


A A C E C R R


A: 2

C: 1

E: 1 (odd so increment numOdd by 1)

numOdd: 2


A A C E C R R


A: 2

C: 2 (even so decrement numOdd by 1)

E: 1

numOdd: 1


A A C E C R R


A: 2

C: 2

E: 1

R: 1 (odd so increment numOdd by 1)

numOdd: 2


A A C E C R R


A: 2

C: 2

E: 1

R: 2 (even so decrement numOdd by 1)

numOdd: 1


By the end, numOdd is only 1, as it should be because we only have one unique character, “E”, that has an odd number of occurrences. And since we end with numOdd <= 1, isPermOfPalindrome will return true. You’ll notice with this logic we must scan every character of the string, as even though a character might have an even count at the beginning, it might eventually have an odd count at the end, and vice versa. We just won’t know until the end. 


Now we’re ready to translate this algorithm over to code. I’ve included my java implementation below:


I’m using a HashMap to keep track of the character counts, and just one integer, numOdd, to track the # of unique characters that have an odd frequency. 


Drop any comments or feedback below, and feel free to share your own implementation coded in another language if you have one!

Thursday, December 31, 2020

Google Foobar - Please Pass the Coded Messages

This week I got around to closing out Level 2 Part 2 of the Google Foobar Challenge - "Please Pass the Coded Messages"

Given a list of digits, our job is to find the order of those digits that yields the maximum number that is divisible by 3.


So a list like [3, 1, 4, 1, 5, 9] would yield 94311, and the list [3, 1, 4, 1] gives us 4311. 


Here’s something to keep in mind that will make our job easier: a number is divisible by 3 if and only if the sum of its digits is also divisible by 3. 


So we could have the following numbers: 375, 753, 357, and 735, and they’d all be divisible by 3, because the sum of their digits is divisible by 3.


If we were to rearrange the digits in the number ‘892’, they will never be divisible by 3, because the sum of their digits will always be 19, and 19 isn’t divisible by 3. We’re left with a non zero remainder each time:


298 / 3 = 99.33333333

829 / 3 = 276.3333333

289 / 3 = 96.33333333

892 / 3 = 297.3333333


This helps a lot. We just need to sort the list once in descending order: so [3, 1, 4, 1, 5, 9] becomes [9, 5, 4, 3, 1, 1], and then we just need to determine the digits to exclude from this sequence while still making it divisible by 3.


Let’s try with our first sequence of digits: [3, 1, 4, 1, 5, 9].


We’ll sort in descending order to yield: [9, 5, 4, 3, 1, 1]. Stringing these digits together we get 954311.


In this case, 954311 is not divisible by 3, so let’s try removing 4 from that series to get us the next largest number: 95311.


Well, 95311 is also not divisible by 3, so let’s take the 5 out and this time try with the same number of digits but keeping the 4 in there: 94311. 94311 is divisible by 3 (since 94311 / 3 = 31437) and so we stop. The answer is 94311.


The next list is easier: [3, 1, 4, 1]. After sorting we get 4311, and 4311 is divisible by 3 (4311 / 3 = 1437), so we end right there.


Pay attention to this tidbit in the instructions : “L will contain anywhere from 1 to 9 digits.”  We’re only going to be given short sequences to work with. Therefore, we can exploit bit masking to represent the combinations of digits that we’ll include/exclude from our number. 


We can start with a bit mask of 1 shifted left by the length of the list (6 in the case of [9, 5, 4, 3, 1, 1]) which yields the binary value 1000000 (or 64 in decimal). We then start counting backwards starting from mask - 1 (63 in decimal or 111111). For each bit position, 1 means we’ll include the digit from that position in the list, and 0 means we don’t include it:


111111 means we include all digits — which would be 954311. 


111110 would be every digit except the last one: 95431.


101010 would yield: 941.


If a combination is divisible by 3 and it’s the max number we’ve encountered, we’ll record it as the max value we have so far. We’ll return whatever that value is at the end. Sometimes it won’t be possible to create any number that is divisible by 3 from a sequence of digits, and in those cases we return 0.


Here’s my full answer in python, which passed all the test cases:


def solution(l):

    lLen = len(l)

    max = 1 << lLen

    l.sort(reverse=True)

    maxAttempt = 0

    for mask in reversed(range(max)):

        attempt = ""

        for index in range(mask.bit_length()):

            attempt += str(l[index]) if (mask >> index) & 1 == 1 else ''

        if attempt == '':

            attempt = 0

        attempt = int(attempt)

        if attempt % 3 == 0:

            if attempt > maxAttempt:

                maxAttempt = attempt


    return maxAttempt


And with that, it’s time to move on to level 3!

Thursday, December 17, 2020

Introduction to Genetic Algorithms

If you are interested in learning genetic algorithms and are looking for a quick introduction - watch this 10 minute tutorial I put together. It introduces you to the basics you'll need to start coding genetic algorithms from scratch. 

Share your questions/feedback/thoughts in the comments!











Monday, November 23, 2020

The 7 programming languages every coder should know

New programming languages seem to emerge constantly in the tech industry. How do you prioritize the ones to learn without spending the rest of your life bouncing from one new language to the next? Regardless of what languages you’re currently coding in, these 7 languages below are must learn for any coder hoping to stay competitive in the job market.


Python - Python is one of the most popular programming languages in the world. It’s straightforward, intuitive syntax, coupled with other aspects like its dynamic typing, contribute to its popularity among not just beginners, but all programmers in general. Its vast offering of libraries has fueled its expansion into a never-ending list of domains: data science, machine learning, statistics and analytics, web development, game making, robotics, are all areas python thrives in. Some of the most popular frameworks (i.e: PyTorch for machine learning, Django for web development, Pygame for making video games) are built on top of Python. Python can be used standalone, or used alongside other languages by extending them with python bindings [include link to bindings page]. Put simply, Python is everywhere. Although it’s certainly not the speediest language, its advantages and pervasiveness wins it a top spot on this list. 


Java - The stance on java seems to be polarized - there are those who love it, and those who avoid it at all costs. Regardless of what group you fall into, the fact is is that it’s in high demand and it’s not going anywhere anytime soon. If you’re looking into a company that is building large-scale enterprise class applications, it’s highly likely their tech stack is java-based. Get familiar with Java, practice building a RESTful service with Spring, and you’re already well on your way to landing a back-end development role. 


C# - A primary rival and competitor to Java, you’ll find this object-oriented language arise in a lot of the same use cases. If a company is producing enterprise applications and isn’t a java shop, there’s a very good chance C# is at the heart of its products. Created by Microsoft, C# was original officially supported by Windows only. Eventually it made its way over to linux and OSX via open source compilers like mono, and now is officially supported cross-platform by .NET core. Anything from little desktop apps to high powered web applications and APIs are written in C#. It will allow you to also break ground into some technologies that are predominantly java territory. Want to build mobile apps for Android, but don’t want anything to do with java? You could always use Xamarin with C# (although I personally wouldn’t recommend doing this based on experiences).


C - You might think its pointless to learn C unless you’re going to build an operating system or write firmware for embedded systems. Why bother with such an “old” technology? Although C is certainly a go-to in low level projects, you should still pick it up even if your plan is to stick with higher level languages. Coding in C will give you a newfound appreciation for the benefits that higher level and managed languages bring to the table. In languages like C# and Java, Garbage collection and memory management, for instance, are all handed in the background while you’re blissfully unaware of the work that is involved in those processes. As such, it’s easy for devs to take these features for granted. Just because you can safely ignore these processes doesn’t mean there aren’t opportunities to optimize your programs from a memory perspective, and coding in C will provide you with the skills to think in ways that will allow you to do exactly that.


Assembly (aka ASM) - When you’re making a function call in C — how are the arguments passed to the function? I don’t mean syntactically: spelling out the function name, followed by parentheses and the comma separated arguments, as is the common method in many languages, but I mean…. how are they actually passed? This is but one example of something so low level that you aren’t typically aware of it — even in C. Just like It’s hard to understand all the things that higher level and managed languages get you until you’ve programmed in C, it’s hard to realize the benefits you get with C until you’ve written in Assembly. Assembly is just about as low-level as you can get, second only to literally writing the machine code instructions yourself. Each architecture has its own assembly language. If you were going to assemble a program for Intel and AMD processors, you’d be writing either x86 or x64, (or both depending on if you’re assembling code for 32 or 64bit). Embedded devices (things like digital cameras, mobile phones, portable game systems) commonly run MIPS and ARM architectures, and so for those you’d be writing MIPS assembly and ARM assembly, respectively. All code, no matter how high level, will eventually result in the actual native machine code executing, so it makes a lot of sense to understand the language that boils down to machine code. 


Javascript - Virtually every modern web page you encounter will have some degree of javascript running in the background. Besides being the backbone of front-end frameworks like React, it can be used for back-end development as well via runtimes such as node.js, PurpleJS, and RingoJS. It’s also arguably the easiest of this list to get started with. You don’t need much to get started with it — you can start coding client-side javascript right now with the default text editors and browsers that ship with your OS. 


SQL - If you code professionally you’re going to come face to face with a SQL query at some point or another. Unless you’re planning on specifically becoming a SQL developer, you don’t need to master every minute detail and construct. However, every developer should be proficient in the basics: SELECT/INSERT/DELETE/CREATE TABLE statements, joins, views, etc. You should be fine with the basics when you’re a dev working in a codebase using and reading SQL in conjunction with the other primary project language(s).


By fostering a deep understanding of Python, C#, and Java, you’ve placed yourself in a high sought out position in the job market. Additionally, you’ll have a strong foundation of the OOP paradigm you can translate over to other object-oriented languages like C++ should you need or want to. Although learning lower level languages such as C and Assembly might not align with your immediate goals, they will give you a deeper understanding of the inner workings of software — instilling you with a mindset focused on efficiency and optimization. You’ll leverage this mindset to make wiser and more optimized design decisions, letting you stand out amongst peers who don’t have access to these experiences. 

Friday, August 28, 2020

Google Foobar: Level 2


I finally had time to start Level 2 of the
Google Foobar Challenge. This level contains 2 challenges. 

The instructions for challenge 2A ("Lovely Lucky LAMBs") are as follows:



At first glance this seems like this is going to require coding in a bunch of rules — but once you extrapolate the patterns from the hints and test cases they give you, the actual code is a breeze to implement. 


The calculations for determining the max and min # of henchmen are going to have separate formulas, so you might want to break this out into two functions — at least initially, to make it a little easier to work with and test the min/max cases. So a template for the solution might look a little like this: 


def min_henchmen(total_lambs): 
   return count


def max_henchmen(total_lambs):

return count

def solution(total_lambs):

return max_henchmen(total_lambs) - min_henchmen(total_lambs)


Let’s address the logic for the ‘stingy’ cases first. When we’re stingy with the LAMB payouts, we’re giving the least amount of lambs to the most amount of henchmen — all the while keeping them all happy and content. As the included test cases states: if you were as stingy as possible with 10 LAMBs, you could pay 4 henchmen with 1, 1, 2, and then 3 LAMBs. This satisfies the requirements because:


- The first most junior henchman receives exactly 1 lamb - (1 lamb like above)

- The 2nd most junior henchman gets at least as many LAMBs as the most junior henchman (1 lamb)

- A henchman will revolt if the share of their next two subordinates combined is > their share (Well, we gave out 2 here which is at least 1+1, so we’re good)


Following this logic then, if we wanted to be as stingy with 12 lambs, we could compensate 5 henchmen: 1, 1, 2, 3,  and then 5 LAMBs.


And so on. Seeing a pattern here? The max # of henchmen that we can give lambs out to can be determined by the Fibonacci sequence. We can handout LAMBs to more henchmen until the sum of all the LAMBs are less than or equal to total_lambs.


Now we just need to work out our min_henchmen() function. The point this time is to be as generous as possible with our LAMB payouts. This means they’ll be fewer henchmen receiving LAMBs, and we’ll have to fulfill all the requirements just like before. 


So again let’s consider our first case. If we have 10 LAMBs in total, fulfill all the requirements to keep the henchmen happy, and be as generous as possible, we’d be able to pay just 3 henchmen: 1, 2, and 4 LAMBs, respectively. In this scenario, we wouldn’t be able to pay a fourth henchman with only 3 remaining lambs without causing a revolt. In order to pay 4 henchman, we would need at least 15 lambs. So the payouts would go like this: 1, 2, 4, 8, etc (respectively). In a nutshell, calculating the minimum amount of henchman (while maximizing the # of LAMBs each henchman receives) is be found by a basic geometrics series.


And then our solution function is easy - we simply return the difference by whatever was returned from the max_henchmen() function - min_henchmen().


This is my complete implementation in python:



I didn’t test the code extensively, but it passed all the test cases. 




On to the second challenge of level 2!

Thursday, August 13, 2020

Google's Secret Foobar Challenge


The other day while googling around I came across this message: 


foobar invitation in a google search



“Curious developers are known to seek interesting problems. Solve one from Google?” Sure, why not?  Upon accepting the challenge I was logged into foobar and presented with a greeting that told me I had managed to infiltrate “Commander Lambda’s” sinister organization.

foobar's greeting message


I did some quick searching around to find other people who had encountered this, and what their experiences were with it. Turns out this has been around for quite some time —but can be quite elusive. Getting into foobar can be tricky, and there’s no guarantee you will. Luck is a major contributor to stumbling upon this, and it’s highly dependent on searching for the right programming related keyword(s). When you’re lucky enough to search for the right topic though, the page will slant a little and the banner with the foobar invite will appear at the top of Google like the one I got.

There are five levels in total. Each level contains a number of challenges. The amount of challenges increases per level along with complexity. There’s some randomness to the content prepared for each foobar invite: your level 1 might not be the same as someone else’s level 1, your level 2 may or may not be the same as someone else’s level 2 content, etc. Each challenge goes like this: you are presented with a puzzle, and you have to code up a solution that will solve that puzzle. Once you start a level, a timer starts counting down and you have until the timer ends to complete and submit a solution. You navigate through the foobar console using basic *nix like commands — if you have any experience with any of these the interface will be pretty intuitive for you. You have the choice of submitting a java or python solution for each challenge.

Preview of foobar ui and code editor


It seems that in its earlier days foobar was utilized as a secret hiring tool for Google to spot talented coders. It seems that it is no longer actively utilized for that purpose, however. Still, it's a fun exercise to try, and even if you don't end up getting a prize out of it, the challenges can help to sharpen your algorithmic thinking skills and get prepare you for any future coding tests and challenges you might face. 

As I get around to it, I'll post my code and walkthroughs for each level of the challenge.


Level 1: Prison Labor Dodgers

This was the first challenge I encountered and the premise of it was pretty simple:

Two lists are given to you, named ‘x’ and ‘y’, that might look like this:

x: [13, 5, 6, 2, 5], y: [5, 2, 5, 13]

Or this: x: [14, 27, 1, 4, 2, 50, 3, 1], y:[2, 4, -4, 3, 1, 1, 14, 27, 50]

In any case, the goal is to identify the unique id, or basically the number that is present in one of those lists but not the other. In the case of the two lists above, the solutions (unique IDs) would be 6, and -4, respectively.

One thing to keep in mind is that there won’t be any inherit ordering or sorting of the numbers in each list, so you can’t count on this in your solution if you intend on taking any shortcuts.

This is what I came up with for a quick & dirty solution:

def solution(x, y):
    first = x
    second = y
    if len(x) > len(y):
        second = x
        first = y
    exists = {}
    for ID in first:
        exists[ID] = True
    for ID in second:
        if ID not in exists.keys():
            return ID

First I make sure the lists are ordered and processed accordingly based on their size (larger list gets scanned first to ensure we look at all possible values). We then iterate through the second list, and check whether that number exists in our dictionary; if not then that’s our unique ID and we return that. 

I decided to roll my own implementation rather than use any of python’s features (you could, for instance, convert both lists to sets and then compute the symmetric difference of them). There are many ways you could go about this. Make sure to verify your code solution for each challenge prior to submitting it! This will indicate whether there are any test cases your code hasn’t passed — so you can go back and try to find what part isn’t working. If you wait to submit a solution only to find out it fails a couple cases, you will fail that challenge. I’ve only done 1 challenge so far and passed all the test cases… so I don’t even know yet what happens if you fail a level.

Another note: Google has test cases prepared in the background - but you don’t actually know what these are, so your solution actually has to work. Expect that google has thought of the various edge cases that could occur from your code, because they typically have, and have built those in to the tests.