Thursday, January 10, 2008

Paragon NTFS for mac update

In the comments of my last blog about the quick hack benchmark I did of paragon NTFS for mac OS X, Anatoly, the product manager from paragon replied. I'm reposting it here so it's not hidden behind that tiny little '1 comments' link at the bottom of the post.
Dear Orion,

My name is Anatoly.
I am Product Manager for Paragon NTFS for Mac OS X driver.



Thank you for your time and efforts to measure the performance of Paragon NTFS for Mac OS X driver.



Frankly speaking your results are not exactly correct for real time usage of the driver.
First of all, the Finder application handles files (copy, create,...) using 2MB block size rather than 512B you tested (the "dd if=//tmp/bigfile of=/dev/null" command uses 512KB block size by default).
Second, to get precise figures you have to unmount/mount partitions every time you perform any test (the reason you got - 87.45MB/Sec). 



So, we retested our driver and would like to show you our results.
We used commands that are similar to yours:



For write:
dd if=/dev/random of=/Volumes/bigfile bs=2m count=100

For read:
dd if=/Volumes/bigfile of=/dev/null bs=2m



HFS+ Firewire: 
Write (MiB/sec) - 4,26; 
Read (MiB/sec) - 36,06.


NTFS Firewire: 
Write (MiB/sec) - 4,24; 
Read (MiB/sec) - 35,26.




Please note in case we will use "bs=1m" we get:


HFS+ Firewire: 
Write (MiB/sec) - 4,34; 
Read (MiB/sec) - 39,29. 


NTFS Firewire: 
Write (MiB/sec) - 4,30; 
Read (MiB/sec) - 42,25. 



According to our tests we can assert that our driver has the same performance as the native HFS+ driver has.


Let me know if I am wrong.



Thank you,
Anatoly.
Well, Wow. I always feel special when important people from companies reply to me! If anyone is looking for numbers, use those ones, as he obviously is far more clued up about it than I am.

I completely agree with his assertion that the driver performs as well as native HFS+

At any rate, I'd already purchased the product, and it's been great. If you are like me and need to access NTFS drives from your mac, you really should buy it.

Thanks!

Sunday, November 25, 2007

5 minute performance picture: Paragon NTFS for Mac OS X

I have a macbook pro, and a large amount of files, and I like to play computer games.
So, I have a large external firewire/USB2 hard drive, and boot camp.
This also means I care about NTFS access from OSX. 
I'd been running MacFuse + NTFS3G. The performance was not toooo bad, but it was chock full of bugs. Drives would show up as network drives, and be called "-n External" and "-n" instead of "External". Not to mention that sometimes stuff would just randomly break. Files would sometimes disappear or move around in the finder and sometimes I just simply couldn't mount the drive. It sucked pretty hard, so I ended up booting up vmware and accessing that drive via vmware's USB2 mapping + samba under the windows VM. Not cool
Suffice to say I was very happy when I saw the release of Paragon NTFS for OSX
I downloaded it, got rid of MacFUSE and NTFS3G, and ran some benchmarks.
Before the benchmark results, let me first say that even if it was just as slow as MacFUSE/NTFS3g, Paragon NTFS would still be worth a look, because it seems (so far) to be rock solid. Drives show up as proper drives in the finder. The volume labels are fine, as is everything else I can see. There is no lag, and I even now have the option of backing up my boot camp partition with Time Machine. Basically it's as if Apple had actually bothered to implement full NTFS support in leopard. That's cool.
Anyway, Benchmarks:
To get the write speeds, I did this:
dd if=/dev/random of=//tmp/bigfile bs=1m count=200
For the read speeds, I did this:
dd if=//tmp/bigfile of=/dev/null
Yes I am aware this is a crap method of benchmarking drives/filesystems. I'm not anandtech and I don't have days to do this.
Computer: MacBook Pro 2.2ghz (the cheapest one)
External NTFS drive: 7200RPM 500gig seagate with 16 meg of cache
External HFS+ drive: 7200RPM 160gig seagate with 8 meg of cache
Both use the identical dirt cheap firewire/USB2 enclosures I found at the local PC shop
  WRITE (Bytes/Sec) WRITE (MB/Sec) Read (Bytes/Sec) Read (MB/Sec)
HFS+ Firewire 6049969 5.77 91697378 87.45
NTFS Firewire 6645725 6.34 19899810 18.98
HFS+ Local 6565372 6.26 90154137 85.98
NTFS Local 6495106 6.19 16776180 16.00
Conclusions:
HFS+ is obviously doing some kind of caching on those reads, as there's no way you can get 85+MB/sec off a plain old 7200rpm drive, let alone the 5400rpm Local drive in the macbookpro. For Actual Use, I can't tell the difference between the NTFS and HFS+ drives
Also, the read/write speeds suck compared to the 30/25 odd MB/sec windows reports when reading/writing files to the disk. But windows lets you enable write caching for removable drives. Maybe OSX doesn't do this. I don't know.
Apart from that, it keeps up with HFS+ and in some cases beats it.
That's Not Half Bad. I might send some my hard-earned paragon's way.

Tuesday, October 30, 2007

How to manually send an email using Rails' ExceptionNotifier Plugin

We have a situation in our rails app where we want to catch an exception and display a custom error message to the user, BUT we still want the exception notifier to fire, so we know all the detailed backtrace data etc, and can deal with it if it's a problem on our end.



Without Further ado, here is the code.

begin

    # b0rk b0rk b0rk

rescue => exception
    fake_params = { :id=>some_id, :etc=>'etc' }
    fake_request = ActionController::AbstractRequest.new
    fake_request.instance_eval do
        @env = { 'HTTP_HOST'=>'fake_host' }
        @parameters = fake_params
    end

    ExceptionNotifier.deliver_exception_notification( exception, ActionController::Base.new, fake_request )

Enjoy :-)



Monday, June 04, 2007

5 Things that I don't like about Ruby

I can't remember the quote or source, but there's a pseudo programmer-interview question which goes something like this: "What's your favourite programming language?" "OK, what are 5 things that are wrong with it that other languages do better?" This is something I've thought about from time to time, and so I figure I'll give it a shot. Obviously ruby is my favourite programming language at the moment, mostly(at the moment) due to the map and inject functions :-)

1. Green Threads are Useless!

The ruby interpreter is co-operative - it can't context switch a thread unless that thread happens to call one of a number of ruby methods. This means that as soon as you hit a long-running C library function, your entire ruby process hangs. I encountered this situation, and tried then to ship it out to another process using DRb. This was even more useless, as when you do that, the parent process blocks and waits for the DRb worker process to return from it's remote function... which doesn't happen as the worker is blocking on your C library function :-( I ended up having to create a database table, insert 'jobs' in it, and have a seperate worker which polled the database once a second. STUPID.

2. You can't yield from within a define_method, or write a proc which accepts a block

It appears to be to do with the scoping of the block, but in ruby 1.8.X, this code doesn't work: class Foo define_method :bar do |f| yield f end end # This line raises "LocalJumpError: no block given", even though there obviously is a block Foo.new.bar(6){ |x| puts x } The other way to skin this cat is as follows, which also doesn't work :-( class Foo define_method :bar do |f, &block| block.call(f) end end # The "define_method :bar do |f, &block|" gives you # parse error, unexpected tAMPER, expecting '|' # :-( This means there is a certain class of cool dynamic method generating stuff you just can't do, due to stupid syntax issues. Boo :-(

3. The standard library is missing a few things

I vote for immediate inclusion of Rails' ActiveSupport sub-project into the rails standard library. I'm sure I won't be alone in thinking this.

4. Some of the standard library ruby classes really suck.

Time, I'm looking at you. Strike 1: Not being able to modify the timezone. Seriously, people need to deal with more than just 'local' and 'utc' timezones. Yes I know there are libraries, but they shouldn't need to exist. Timezones are not a new high-tech feature! Strike 2: The methods utc and getutc should be utc! and utc, in keeping with the rest of the language. This alone has caused several nasty and hard-to-spot bugs Strike 3: What the heck is up with the Time class vs the DateTime vs the Date class? This stuff should all be rolled into one and simplified. The Tempfile class is also notably annoying. Why doesn't it just subclass IO like any sane person would expect?

5. The RDoc table of contents annoys me

This is probably more "Firefox should have a 'search in current frame'" feature, but under http://ruby-doc.org/core/, have you ever had the page for say Array open, and wanted to jump to say it's hash method? I usually do this using firefox's find-as-you-type, but seriously, try doing just this in the rdoc generated pages with the 3 frames containing everymethodever open. Cry :-(

Monday, April 23, 2007

HOWTO: Create a GParted LiveUSB which actually works WITHOUT LINUX

EDIT:

Turns out there is a windows version of syslinux, to be found HERE.

If I'd kept reading for about 2 more minutes I would have found that out and managed to avoid pretty much all of the timewasting I did last night. Sigh. At least the other people trying to make it work using loadlin indicates I can't have been the only one to get it wrong :-(

Also, the graphics card thing is a non-problem. Just chose Mini X-vesa in the gparted boot menus and it's fine

Moral of the story? Just because you've found a solution doesn't mean it's the best one. Keep looking until you can be sure it is!


So, I wanted to repartition my hard drive tonight. I've used GParted before and it was brilliant, so off I went to download the liveCD again.

Once at that site, I saw the LiveUSB option from the left-hand menu, and thought "Brilliant, I don't have to waste a CDR and it will be much quicker anyway!"... Little did I know that PAIN and DESPAIR awaited me. I'll publish how I resolved this in the hope that less other people will have to.

Step 1: Download the GParted LiveUSB distro

I clicked 'Downloads', from the navigation, followed the liveUSB links, and wound up here:
http://sourceforge.net/project/showfiles.php?group_id=115843&package_id=195292
I downloaded gparted-liveusb-0.3.1-1.zip, and unzipped it. Hooray, now what?

Problem 1: The GParted LiveUSB documentation is crap!

The GParted LiveUSB information here says firstly I need to download a shell script, then I run it and copy some files to my USB key... Apart from a link to one forum post here that's it. Documentation? Instructions? Why do we need those? What could POSSIBLY go wrong?

Problem 2: Running shell scripts on windows doesn't work too well

The above shellscript invokes syslinux, and just about everything else on the net that talks about creating bootable floppies/USB keys also sooner or later invokes syslinux also. This seems to set up the boot record on the USB key so that you can boot linux off it. DOS used to have a utility like this called 'system' or 'sys' or somesuch but I can't remember. Seems simple enough, except I NEED LINUX TO RUN IT. Actually no I don't... see above. oops

In my humble opinion, if I was running linux already, I wouldn't need the liveUSB, I'd just apt-get install gparted and run the damn thing. Yes some travelling sysadmins might have a linux box at home and also need a usb key to take around, but I'm not one of them. The entire reason I'm trying to get this liveUSB to run is because I DON'T have linux.

So, I read that forum post, and noticed at the bottom someone using loadlin to load linux from a DOS system. Aha!

Step 2: A whole crapload of google searching and researching...

As I can't make my USB key linux-bootable without linux, I need to make it DOS-bootable, then get loadlin to load the linux kernel that comes with the gparted liveUSB. I'm going to skip all the boring details as it took me frickin ages and just explain what to do...

Step 2.1: Download a DOS bootdisk so we have DOS

Goto http://www.bootdisk.com/bootdisk.htm and download the "Windows 98 SE Custom, No Ramdrive" boot disk. This gets you an executable which expects to write to your floppy drive... except I don't have a floppy drive. BAH.

Step 2.2: Extract the DOS bootdisk image with WinImage

  • Goto http://www.winimage.com/download.htm. I went for "winima80.zip" as I just wanted to run it once without the installer guff.
  • Run winimage. Do File->Open, and point it at the boot98sc.exe file you downloaded in step 1.
  • Once this is open, chose Image->Extract, and dump all the DOS system files somewhere
  • Step 2.3: Make your thumbdrive bootable

  • Goto http://h18000.www1.hp.com/support/files/serveroptions/us/download/20306.html and download the HP Drive Key format utility. As far as I can tell this is the easiest way to make your USB key bootable. It works with pretty much everything not just HP keys.
  • Make sure your USB key is plugged in
  • Run the HP program, and format your USB key using FAT (FAT32 should work too, but I didn't try it). Make sure to select "Create a DOS startup disk", and in the "using DOS system files located at:" box, enter the directory you dumped the DOS system files from winImage earlier
  • Hit start, and wait for it to finish. JUST IN CASE YOU FORGOT, THIS WILL ERASE ALL THE FILES ON YOUR USB KEY, SO BACK THEM UP FIRST, K
  • Step 2.4: Get loadlin

  • Goto http://distro.ibiblio.org/pub/linux/distributions/startcom/DL-3.0.0/os/i386/dosutils/ and download "loadlin.exe" to somewhere on your PC
  • Step 2.5: Copy files onto your USB key

  • Unzip "gparted-liveusb-0.3.1-1.zip" if you haven't already, and copy all the files into the root of your USB key. Your USB key should now contain those files, COMMAND.COM, IO.SYS, MSDOS.SYS and nothing else. No directories etc.
  • Also copy loadlin.exe into the root of your USB key
  • Step 2.6: Make loadlin run automatically

    Note: This is like in the forum post here, except it actually works. I think that's out of date.
  • In the root of your USB key, create a new file called "loadlin.par"
  • Open it with notepad or something, and put this in it: linux noapic initrd=initrd.gz root=/dev/ram0 init=/linuxrc ramdisk_size=65000 (for those interested, those are the kernel boot parameters which I stole that out of syslinux.cfg from the gparted liveUSB distro. If that file changes, so should your loadlin parameters)
  • In the root of your USB key, create a new file called "autoexec.bat"
  • Open it with notepad or something, and put this in it: loadlin.exe @loadlin.par
  • Step 3: GO GO GO

    Reboot your computer! If you've set up your BIOS properly to boot off USB keys, your computer should now boot the GParted liveUSB. HOORAYZ!!!!1111

    Step 4: cry

    That's as far as I got, because the version of X.org on the liveUSB doesn't seem to like my NVidia 7600GT, so I'm stuck with a command prompt. Those of you with other graphics cards however should be fine. Whether the liveDistro includes command line partitioning tools I dunno, I might go look at that now.

    If anyone would like to copy/distribute these instructions, or edit copies/etc, you are free to, as I am putting this particular blog post in the public domain under the creative commons public domain license.

    Sunday, March 04, 2007

    Rails 1.2 changes

    Formats and respond_to

    In earlier versions of rails you could have your actions behave differently depending on what content type the web browser was expecting – eg:

    respond_to do |format|

        format.xml{ render :xml=>image.to_xml }

        format.jpg{ self.image }

    end

    However to make this work you needed to set the HTTP Accept header in the HTTP web request. This is hard to do outside of tests. A new default route has now been added

    map.connect ':controller/:action/:id.:format'

    The additional format parameter lets you override the format so you can now load people/12345.xml or image/12345.jpg in your web browser to test what happens instead of mucking about with HTTP headers.

    Note you still have to register MIME types for the formats you need – for format.jpg I had to put

    Mime::Type.register 'image/jpeg', :jpg

    In my environment.rb, as jpg is not noticed by default

    Named Routes

    map.index '/', :controller=>'home', :action=>'index'

    map.home '/:action/:id', :controller=>'home'

    These create a bunch of helper methods which you can use anywhere you'd supply a URL or parameters for a redirect – eg:

    def first_action

        redirect_to index_url # redirects to /

    end



    def second_action

        redirect_to home_url( :action=>'second' ) # redirects to /second

        # which is the 'home' controller.

    end





    <%= link_to 'home', index_url %>

    <%= link_to 'test', home_path( :action=>'test' ) %>

     

    The difference between foo_url and foo_path is that foo_url gives the entire url eg: http://www.site.com/people/12345 whereas foo_path just gives /people/12345

    Gives your code lots more meaning and makes it shorter. Definite win for commonly used things.

    Resources

    CRUD means Create, Read, Update, Delete.

    These map to the four HTTP methods – POST, GET, PUT, DELETE.

    HTTP methods let you have shortcuts, so instead of /people/create you can just do an HTTP POST to /people. Also /people/show/1 maps to GET /people/1, etc etc

    Routes are created differently – for the above it is

    map.resources :people.

    Run script/generate scaffold_resource people to have a look

    NOTE: Rails expects resources in both the routes.rb and controller names to be named in plural - eg:

    www.example.com/people/1 instead of www.example.com/people/1

    Philosophy

    Basically they are encouraging you to write your controllers and app so that everything revolves around either a create, read, update, or delete of some resource.

    Contrived Example: User Login sessions:

    Old way – Revolves around action:

    User Logs in – post a form to /users/login – this sticks a 'Login Token' of some sort in the session to identify them.

    User does stuff – look up the session and link it back – might put an is_logged_in? method on your user model or something.

    User Logs out – posts a form to /users/logout – this removes thing from the session.

    New way – Revolves around resources

    Identify what the 'resource' is – in this case it's the Login Token.

    User Logs in – Create a LoginToken by POSTing a form to /LoginTokens – stick it's id in the session or something

    User does stuff – Find the correct LoginToken based on it's id, check it's valid etc.

    User Logs out – Delete the LoginToken by DELETEing /LoginTokens/1

    Conflict of interest?

    This LoginToken is behaving a lot like a model even though it's a controller. In fact you should create a model for it. The LoginTokensController should only be a lightweight wrapper around this model. This is a definite win if you can structure your app like this because it seperates the different areas of code out.

    In the old way we had the users controller handling login, logout, and whatever else it needed to do – probably about half a dozen other unrelated things. This gets you messy code which is hard to understand/modify. By moving each part out to its own controller we end up with several separate nice clean controllers instead of one big messy one – easier to maintain and to see who's responsible for what. Very important!

    REST API's

    Many blog entries talk about how you can get an externally accessible API 'for free' by extending these CRUD controllers. The standard example is something like:

    Now that your users are accessible via GET,POST,etc to /users/1, we can extend that controller using respond_to so that you can also query it for an XML or JSON representation of the user – this can then be used by other websites/apps for free! Hooray!

    This is nice in theory but not so nice in practice. Why?

    If you are in a webapp, doing a POST to create a new LoginSession will result in a redirect_to home_url or something like that. However for an external API, you're meant to return an http response code of 201 – Created to indicate the create was successful. Trying to jam these 2 things into the same controller is a mess.

    This does not mean the REST idea is bad, only that you need to think a bit more.

    If you've done the right thing and created a LoginSession model, then you can just create 2 lightweight controllers – one which fits into your web app, and another if you like which processes XML/JSON.

    You still get the major benefit which is that by thinking of stuff as resources, you get a much better design/structure of your app.

    ActiveResource

    If you have an external XML REST api, these resources end up looking a lot like some data that you might want to load, update, store, etc, like a kind of remote database.

    They therefore decided to make something called ActiveResource which would do for XML REST resources what ActiveRecord does for databases (in a limited fashion)

    For Example:

    class RemotePerson < ActiveResource::Base

        set_site http://localhost/rest_people ## the base URI

        set_credentials :username => "some_user", :password => "abcdefg" ## for HTTP basic authentication

        set_object_name 'person' ## the other end will expect data called 'person' not 'RemotePerson'

    end

    You can then do

    RemotePerson.find( 1 )

    This will fire an HTTP request at http://localhost/rest_people/1. It will load the resulting XML and convert it into an object. You can change its data, and call save, etc like you would with a piece of data from the database. When you have 2 sites that need to communicate with each other, this makes it a WHOLE lot easier

    The bad news – This isn't in rails 1.2 They pulled it out in one of the beta versions and there's no indication as to when it's coming back

    The good news – I wrote one (a limited version thereof anyway) to replace what didn't ship with rails.

    to_xml, from_xml, to_json, etc

    For all these active resource things to work, they need an easy way to convert data to and from XML so it can be sent over the HTTP request. There are now new methods - to_xml, from_xml, to_json, and other stuff like that which will convert the object to and from xml. These have been added to Hash, ActiveRecord, and other things like that.

    Multibyte Support

    Rails 1.2 hacks the string class so it is now sort of Unicode (UTF-8) aware.

    TextHelper#truncate/excerpt and String#at/from/to/first/last will now automatically work with Unicode strings, but for everything else you need to call .chars or you will break the string.

    In other words, If you need to deal with foreign characters (and for the US we probably will) String.length is broken and so is string[x..y] or just about anything else you'd want to do! Beware!

    I have no idea how this is going to impact storing strings in mysql etc.

    Tons of other bits and pieces

    Lots of things have now been deprecated doing stuff like referencing @params instead of params etc etc – these all get dumped to the logs as deprecation warnings – they are still ok now but will break in rails 2.0

    image_tag without a file extension is also deprecated. You should always specify one.

     

    See here:

    http://www.railtie.net/articles/2006/09/05/rush-to-rails-1-2-adds-tons-of-new-features

    Friday, November 10, 2006

    Why everyone wants to get rid of the parentheses in lisp

    This post is pretty much a response to http://eli.thegreenplace.net/2006/11/10/the-parentheses-of-lisp/ It's a good post - if you haven't read it, eli seems to say that: He's noticed a lot of people trying to use whitespace to remove most of the parens in lisp, and can't understand why. His opinion (as it seemed to me) was that removing them would be counterproductive because the parens (and their uniform syntax), which is what makes lisp so much better than everything else).

    My 2c:

    The case FOR s-expressions

    * Uniform syntax is theoretically very appealing (from a purist point of view)

    * It lets you write macros. Macros are incredibly powerful and basically awesome. I have macro-envy in most languages most of the time.

    The case against s-expressions

    * s-expressions are not at all like how I think about things (or how anyone else who is not a die-hard lisper thinks about things), because the lack of syntax is completely at odds with natural language.

    To elaborate on that last point: Because english is my native language, that's how I think. A fair amount of the time, code which I consider "good" ends up looking (and structured) like a shorthand version of english, because that's what I find is clear and understandable.

    I can write Ruby, C++ and C# and most other blub languages in such a way that maps relatively closely with how I think. In short, they 'fit my brain'. I also realise that over time my brain has adjusted to fit them as well, so I am aware that this kind of thing is probably always going to be biased towards the incumbents.

    Conclusion

    I believe that with the exception of a few idealists who hold the 'uniform syntax for purity's sake' argument above all else, pretty much everyone else is in it for the macros. To get them, I see two paths -

    1) Change yourself: Do enough lisp programming, for a long enough time, and put in the effort to make your thought process (at least with respect to programming) match with s-expressions. This seems to be what all the 'hardcore' lispers have done, and the evidence seems to point towards this having a huge and amazing payoff. I'm nowhere near this goal, but this is eventually what I'm aiming for with my casual lisp programming... It is however, (like any learning) long and dare I say, hard.

    2) Try and change lisp: Try and make the syntax fit better with your brain, so you (and all the other regular programmers) don't have to put in the hard yards adjusting quite so much. This is what I believe the lisp-whitespace people are aiming for.

    I do find the whitespace-lisp easier to read/understand than normal lisp, and it doesn't seem to sacrifice any functionality (macros still work), so it could be a winner.

    Then again, significant whitespace brings in a whole host of other problems, and makes it not so "pure", so perhaps not.

    Is it a good idea? In the end, I vote for a definite "maybe" :-)

    Tuesday, October 17, 2006

    Programming best practices

    I was going to call this "X secrets of highly effective developers", like some other people, only these things shouldn't be secrets. Note this is as always my not-so-humble opinion, so it is entirely likely that this article is either a) misguided, b) missing things, or c) entirely wrong, but I can't give you anyone else's opinion now can I.
    These are all typical cliché's, but I'd like to try explain them just as a brain dump anyway.

    1. Code for people, not computers.

    This is really the absolute number one goal. Everything else can be taken as a corollary to this.

    If you don't code for people, you are writing un-maintainable code. However, it's easy to throw the term around like a lot of other slogans/buzzwords without actually having a solid understanding. What this means to me is in fact "Try and write your programs as if they were plain english"

    Well, you know, not quite english, because english has it's own giant set of problems too, but the point I'm trying to make is that someone else should be able to read your code and it should flow as if it has topics, headings, sentences, paragraphs, and so on. You should read it like a book, not decipher it like a code. In fact, code is a crap word, we should call it something like "instructions", instead of code.

    How do we do this? With years of experience, and a constant drive to learn and apply new things.
    But for now, here's a couple of pointers which I find have helped me so much that I feel like slapping other developers who don't do them:

    2. Use good names

    This is obviously a subset of #1, but if I had to pick the most important thing, this would be it. Again, "Use good names" is a meaningless phrase, what you should do is "call things what they are". Whenever you have a variable/function/class/whatever, ask yourself "what actually is it? a buffer? a file handle? a person? what?", and then call the variable that. Simple, but mostly always overlooked by most programmers. Other developers should almost always be able to look at a variable/function/class and make a correct guess as to what it is, and what it might do otherwise you probably have a bad name.

    This is doubly useful because sometimes you will have trouble expressing what a something actually is. Sometimes, this is just not being able to think of the right word (go learn english :-D ), but more often than not, it is a strong signal that your design isn't correct, or you have accidentally gone down the twisty path towards a tangled mess of garbage.
    For example, if you can't think of a single good name for a class or function, it's probably because it's doing more than one job, and should be broken up into 2 classes/functions.

    At the end of the day, if you can't even think of a reasonable name for a thing which makes your program work, then how will anyone else (usually you, 6 months later) ever hope to understand it?

    3. Use abstraction

    This can be "bottom up" programming, or "top down" or whatever design methodology is in fashion at the time, but the important thing is that you build code On top of other code, not alongside it. You are creating a pyramid, not tiling a floor.
    That may not have made too much sense - to try explain it a bit better, think of writing a program that opens a file, writes some string to it, and closes the file.

    If you were to use the all-too-common floor tiling method, you'd have a big long function (or lots of small functions running in sequence, whatever), which would do the following in sequence:
    Allocate a file handle -> call the API open function -> allocate the memory for the string -> keep track of how many bytes we have -> write the memory to the file using the API file writing function -> close the file -> free the string memory.
    All the operations are on the same "level", the stuff is just happening in a big row, like laying down tiles next to each other..

    If we are to write the above using abstractions, we'd instead have a file class, and a string class. The file class would deal with the file API (handles and stuff), and the string class would deal with the string memory. Then, instead of our program allocating handles and memory, it can just deal with files and string classes. It would do the following.
    Create a file object -> Create a string object -> call file.write( string ) -> cleanup done automatically by objects.

    When you write programs correctly with abstraction, you can stack the abstractions on top of each other, leading eventually to code that is almost like pseudocode or a domain-specific-language.
    This is what object oriented programming, and a lot of other techniques are actually for (as opposed to the common retarded view of inexperienced programmers, that OO isn't being done correctly unless all your classes use inheritance somehow)

    4. Do the simplest thing that can possibly work

    Now before I get branded as an agile zealot, not everything from agile is actually bad. What this actually means is Do the right thing, but do it the simplest and smallest way you can. Don't write code which doesn't directly help you get things done, or tries to solve problems you don't actually have.
    This also doesn't just apply to your higher-level design, but low level too. Functions/classes/interfaces/etc should all be as simple as possible, and do only what they need to.

    A classic example of doing the wrong thing here is building a big pile of classes and interfaces and message-handling code before you actually attack your main problem. Yes it's fun to build frameworks and architecture, but at the end of the day, it'll probably just get in the way.

    5. Don't repeat yourself, refactor instead.

    If you are like lots of developers I've seen, and believe the best way to start a new program/library/class is to find a similar one and copy/paste it into a new file, STOP NOW. BAD PROGRAMMER! SIT!

    If you find yourself writing the exact same code twice, refactor it into a common function or class.

    If you find yourself writing similar code twice, refactor the common bits into another function or class (generics, dynamic types or other ways of dodging verbose typing, and first-class functions are a big win here), and have the remaining different bits as small and clear as possible.

    5. Good design does not come from 'design,' but from refactoring.

    If you do all the other stuff, you'll probably find yourself left with a ton of small functions and classes, and all your other classes will be using them. This is already better than lots of giant functions which duplicate code and functionality, but can get a bit messy provided you don't clean them up.
    Most likely however, a bunch of your helper functions will all take similar parameters. These are prime candidates for making a new class. Remember also, not everything should to be a class. It's fine to have a bunch of global functions in a namespace (or a static class if you're stuck in C# or java), if that's the nicest way to think about those kinds of objects. The main goal is to always try and make sure that your helper/library functions are as simple, clean, and useful as possible, and are arranged in clear and obvious groups. You can even write unit tests for them :-)

    Sooner or later, you'll probably end up with either some kind of framework (to my knowledge, this is how Rails came about), or a set of general classes, like the .NET base class library, but for dealing with the kinds of problems your company faces, or perhaps both.
    This is excellent, as you'll be able to re-use these things over and over again in future, meaning next time you have to write that stupid app which has to read the registry and write to files, you'll be able to use your framework/library code and do it in 5 minutes instead of a day. Your boss will love you, and you won't mind programming in C++ anymore because you can actually get things done now.

    Also, because you'll have created this framework/library code based on other working code, and based on what you actually need to do, and refactored it to best fit your problems as you go, you'll have stuff which is actually useful and good.
    People are stupid, and 99% of us can't design our way out of a paper bag. Most frameworks that get 'designed' up front wind up completely missing the point and to solving the wrong problems in the wrong way. But, by keeping existing code simple, clean, non-repeating, and constantly refactoring it, we can end up with some well structured and maintainable code anyway. Owzat? :-)

    Sunday, October 15, 2006

    Diffing itunes music libraries with ironpython

    Intro:

    Ages ago when I first played with python, I found it was pretty cool, in a weird sort of way, and that I could probably love it once I had spent 6 months bending my brain around it's quirky bits (too many underscores, whitespace importance, inconsistent and strange libraries)... And then soon after, I found ruby, which had none of the above problems, so I didn't bother learning any more python...

    Until a few weeks ago, when IronPython was released.

    For the uninitiated, IronPython is python which runs on and inside the .NET CLR (or mono, just not as quickly). My biggest blocker for regular python was the built in libraries/docs. I'm bound to be flamed, but the ones I looked at (file access, networking, HTTP, etc) just were not intuitive. IronPython however is a dream, because it uses all the .NET libraries instead (or as well as, if you like). I'm pretty familiar with the .NET BCL, so this was great.

    Actual content :-)

    During my ~4 years at my current job, I have built up a large collection of music which I'd listen to. I also have a large collection of music at home. As I'm leaving in 3 weeks, and will lose all that data, I wanted to take a copy of the music from my work computer home.

    The problem with this is that I only have a 6 gig ipod mini to transfer the songs on, so I can't just copy them all. I needed to diff the 2 music libraries, and only copy the songs that I don't have at home already. iTunes exports a large hairy pile of crap XML file when you ask it to export it's library, so here I thought would be an opportunity to play about with some IronPython, and post it on the net in case it's useful to anyone else.

    Here it is, the comments are the documentation :-). Hopefully it's useful, if only as a quick demo of how things work in ironpython.

    # Import all the libraries we'll need import clr clr.AddReferenceByPartialName( "System.Xml" ) from System import * from System.IO import * from System.Xml import * # Create a helper function to convert the iTunes XML file into a hash so it's actually useful # An example of one of the hash entries: ret[ 'Disturbed: Prayer' ] = 'file://C:/path/disturbed_prayer.mp3' def fileToHash( fileName ): ret = {} doc = XmlDocument() doc.Load( fileName ) # The XPath is ugly... Export your itunes library and take a look to see why for elem in doc.DocumentElement.SelectNodes( "/plist/dict/dict/dict" ): # song name is always the first <string> song = elem.SelectSingleNode( "string[1]/text()").Value # artist is always the second <string> artist = elem.SelectSingleNode( "string[2]/text()").Value # Unfortunately the filename is not fixed in the structure, so we have to # find <key>Location</key> and then move to the next element after it path = elem.SelectSingleNode("key[text() = 'Location']").NextSibling.FirstChild.Value # Add it to our return-param ret[ "%s: %s" % ( artist, song ) ] = path return ret # Parse both files into seperate hashes homeSongs = fileToHash( "itunes library home.xml" ) workSongs = fileToHash( "itunes library work.xml" ) # Create a new dict containing only songs that are at work but NOT at home diffSongs = dict( [ (song,workSongs[song]) for song in workSongs if not homeSongs.ContainsKey(song) ] ) # Write them all to the output file # The format of the output file is: file://C:/path/disturbed_prayer.mp3 # Disturbed: Prayer # The reason I've done it like this with the path first is hopefully to make it easier # for another script to be able to use it as a list of filenames to copy... writer = StreamWriter( "diffSongs.txt" ) for str in [ "%s # %s" % (diffSongs[song], song) for song in diffSongs ]: writer.WriteLine( str )

    Disclaimer:
    1) This was meant to be quick to write, I didn't care about run performance or any other nifty tricks - it only takes 3 seconds to parse 2x 3.5 meg XML files, and that's good enough for me.
    2) Sorry about no syntax hilighting, I couldn't find any decent way to do it short of screenshotting my PSPad window.

    Wednesday, October 11, 2006

    Vista RC1 and RC2: Impressions

    Over the last couple of days I've installed both vista RC1 and RC2 at home, and today installed vista RC2 at work. Here's a quick brain dump of my thoughts about various things that have cropped up: Except for where I explicitly say so below, RC1 and RC2 are pretty much the same.

    Aero Glass (fancy graphics):

    My home machine is an Athlon 64 3200+, with 512 RAM, and a Radeon 9700 graphics card w/128Mb. The graphics card is the main thing that gets hit by aero, and in RC1, aero wasn't enabled by default after the install. A quick look at the control panel to turn it on, and it was fine. I really liked it. Sure, it's just eye-candy, but who said computers have to be ugly? All in all I was very happy....

    HOWEVER: RC2 decides that I need 1 gig of RAM to enable aero. I haven't been able to find any overrides as of yet. This is stupid. Someone at the Microsoft marketing department has gotten their nose into this or something, because I KNOW aero runs fine on this machine with 512 RAM, having just run it the day before. W T F.

    Minor gripe: The "flip 3d" thing they have is useless. It offers less usefulness than just the standard alt-tab. To add to the blogosphere whining, why didn't they just verbatim rip Expose from the mac?

    The Vista Basic Theme (low rent graphics):

    You could just download a theme from deviantart or elsewhere for windows XP, and you quite literally wouldn't be able to tell the difference. It's debatable as to if this theme is uglier than the default XP theme, it's certainly not much  better that's for sure.

    Performance (superfetch smart memory management and caching)

    This is a bit strange, but overall I was suitably impressed. Example: When I'd play Warcraft III on windows XP, quitting would bring a 30s to 2 minute "crunch", while everything got paged around the place. On vista, with identical hardware, it's much more responsive

    Think of it like this:
    In XP, when you load a large program, it gets loaded into ram, then 100% of your CPU/HDD resources are free for use. When you quit, or do some other memory intensive thing, the system takes a dump for a while to sort all itself out.
    In Vista, when you load a large program, it gets loaded, but only 95% of your CPU/HDD resources are free, the other 5% are used by superfetch tinkering away in the background - which means that when you quit (or other memory intensive thing), all the stuff it's done in the background pays off and your system just runs a bit slower instead of just falling over.

    Programs which you never run take about the same time to load, and basically everything performs the same. Frequently used stuff like firefox loads MUCH more quickly, even though...

    Memory usage (oh noes! those horrible background tasks!)

    From (my admittedly fuzzy) memory, a vanilla install of XP RTM would use about 90 meg of RAM. Post installing SP2, it would use about 180. After I did the customary beatdown of all the unnecessary services, it'd use about 120 megs.

    Vista seemed to use about 300 megs of ram post install. After the services beatdown, it got down to 200 meg (not counting disk cache of course). Aero adds about 50 meg to this. However, the system overall seemed just as responsive as XP did - This is basically a testament to how good superfetch is, but vista still whores teh RAM.  Hopefully it won't be quite so bad when they RTM it, but I'm not holding my breath.

    As an aside, this is probably why they don't let you run aero on a machine with 512 RAM - 300 for vista + 50 (at least) for aero doesn't leave much for applications. I can see the reasoning behind it but how about instead of disabling aero on low-ram systems, however there definitely needs to be an "I am not a retard" switch, so that I can use the free ram I got by turning off all the useless background crap to run aero. As I said earlier, I KNOW aero runs fine on this PC. WTF.

    Readyboost

    I didn't test this on RC1 because I didn't have the memory stick then, but on RC2 I am using an entire 1GB usb2 memory stick for Readyboost. I can't provide any actual numbers or anything, but from my very limited experience, it seems to make a noticeable difference in responsiveness. I'm very happy with it, and I'm usually picky about these kind of things.

    PS: Note I said "responsiveness", not "performance". Stuff still runs the same, just those "Crunch" moments like when you load firefox or alt tab out of a game or other "beat the crap out of the pagefile/disk" moments are a whole heap better.

    PPS: No, adding 1 gig of readyboost to a 512 MB system still doesn't let you run aero. bastards.

    Other neat things

    I had mp3's playing, and upgraded my graphics drivers without rebooting or even skipping a beat in the mp3s. It thrashed for a while, the screen went black for a second or 2, and presto new graphics drivers. Seriously impressed.

    Being able to type arbitrary commands, like "net stop server" into the search box in the start menu, and into the explorer address bar is awesome

    Windows explorer and media player now use the same format for album art as itunes, which is sweet.

    The new task manager/performance stuff is great.

    And the winner is!

    Overall, I'd have to say my favourite part of vista overall thus far would have to be the new windows explorer. I like the new clickable address bar format. The revamped 'documents and settings' thing is so much nicer. The customizable 'favourite links' panel is great. The searching and filtering is brilliant. The new start menu is insanely good.
    Oh and not to mention the facts that a) it doesn't hang when you try access network shares, and b) if you're copying/moving/deleting a bunch of files, and one of them fails, it carries on with the rest instead of just falling over like a useless pile of crap like XP and everything before it did. I could go on for hours. I love it. A million points to the shell team at MS.

    And the loser is!

    This is so cliché I know, but user account control sucks. I agree with the principle in theory, but it's implementation just seems to suck.

    On the one hand you have some of the things like when you copy a file to a "restricted" folder - you should get one popup asking for confirmation, but you also before that get a dialog warning you that if you continue you will be prompted for confirmation. They could quite literally rewrite it to the following text: "If you try and do this, we will annoy you with another dialog after this one, are you sure you want us to pop the second dialog so you can click yes and be annoyed"

    On the other hand you have things where it just doesn't kick in. If I use explorer to copy and paste a file to a "restricted" folder, it pops me for confirmation and succeeds. If however I drag/drop the file, it just fails with 'access denied' and no prompt or way to get around it. It seems to have no awareness of things other than explorer in some situations too - I can't save files from firefox to some places; doing things from the command prompt pretty much just doesn't work, etc.

    IMHO, it looks as if they haven't actually implemented UAC as part of the windows OS or API, they've just made administrators into restricted users, and made explorer dick around with permissions when it launches applications. My recommendation? Turn it off like everyone else, and wait another 5 years, maybe MS will get it right next time.

    Conclusion

    Vista overall is worth upgrading from XP. I don't know how much I'd pay for it, but it definitely is an overall improvement and there doesn't really seem to be many other downsides apart from the odd piece of software here and there. Almost everything runs just fine, and it's rock solid stable. Just remember to turn off UAC :-)

    Monday, September 25, 2006

    Function Pointers in C/C++ and boost::bind

    In a previous blog entry, I showed the Windows QueueUserAPC function, and how you could use it to get other threads to execute functions. Now this was kind of cool, but if the only functions we can use are free functions which have only a single 32 bit parameter, not so useful. I said I'd explain how to use boost::bind to solve this.

    Now, I'm not going to explain how this interacts with QueueUserAPC just yet, because explaining boost functions and bind is well big enough for a blog entry of it's own. Here it is.

    Intro - Function Pointers in C and C++

    If you're not familiar with what a function pointer even is, well:

    • All your code that you write gets turned into a big bunch of binary stuff by your compiler.
    • In order for this code to run, it has to load this binary stuff into memory
    • Once the binary stuff is in memory, the CPU can be told to execute arbitrary bits of it. This is what C/C++ does behind the scenes when you normally call a function
    • A function pointer in C/C++ is a pointer to a bit of that memory where a function lives.
    • When more code is loaded, or passed around from one part to another, you can get pointers to this new code as well as the static stuff which you wrote upfront.
    • This lets you do things like call functions which didn't exist when you compiled the program, (ie: DLL's), or tell some code to execute a runtime-specified piece of other code on some event (ie: callback functions)

    If you're not familiar with function pointers in C, here's an example:

    void Print( int a ) {
        std::cout << "Free Function, a = " << a << std::endl; 
    }  
    
    void main() {
         void( *x )(int) = &a_function;
          // we now have an object x, which is a pointer to function of type void(*)(int).
         // it happens to be pointing at our Print function
         x( 5 ); //call it 
    }


    This syntax is fine and dandy for C functions, which don't have classes or anything, so the types are all pretty simple, however in C++, we have classes, which have member functions (or methods if you prefer to call them that). Imagine the following class:

    class FooClass {
    public:
        void Print( int a ) {
             std::cout << "A FooClass, param = "<< a <<" this = " << this << std::endl;
         }
    };

    Now, if we want to get a pointer to the Print function, we have to write some extra stuff so the compiler can tell that it's a member function of FooClass, and we also have to pass the instance, so the compiler knows which FooClass the Print function should belong to. We might have 50 of FooClass in an array and it's got to be able to figure out which one is the right one.

    FooClass* myFoo = new FooClass(); //create an instance of our FooClass
    void( FooClass::* x )(int) = &FooClass::Print 
    // we now have an object called x, which is a pointer to function of type void(FooClass::*)(int)
    // it happens to be pointing at our FooClass::Print function, but it doesn't know which FooClass instance yet
    
    (myFoo->*x)( 5 );
    //call the function, telling it that the FooClass represented by myFoo is the one to use 
    //, as if we'd called myFoo->Print( 5 );

    As you may well have noticed, using pointers to class members sucks. I've done a fair bit of this kind of thing and I had to go and look up the documentation to remember quite how it was supposed to work. C++ cops a lot of flak for this kind of stuff, and deservedly so in my humble opinion.

    We could always make it nicer by using some typedef's, but it's still ugly, still probably wouldn't make sense to most novice programmers, and we still have the problem that the function pointer doesn't stand alone - we also need to pass the object instance around.

    If they didn't know any better, but wanted to use this kind of thing, most programmers would probably end up creating some kind of struct which contained the instance, the function pointer, and some other arbitrary data, and passing that around the place. That at least makes it possible, but it still sucks.

    boost::function

    If you take the following 3 things:

    • Function pointer type declarations suck
    • You can do lots of amazing stuff with templates in C++
    • The guys who write the boost libraries are really really smart

    Then, for free functions, we get this:
    void Print( int a ) {
        std::cout << "A Free function, param = "<< a << std::endl;
    };
    
    void main() {
        void( *oldFunc )(int) = &Print; //C style function pointer
        oldFunc( 5 );
    
        boost::function<void(int)> newFunc = &Print; //boost function
        newFunc( 5 ); 
    } 

    It's just my opinion of course, but I think the C style function pointer with the variable name in between the void and the (int) is confusing, whereas the boost function just behaves how you'd expect it to. Code that clearly states what it does, and then simply does it, is the best code.

    Even though this is a trivial example and the boost one isn't actually much nicer, I'd still use it just because it's easier to read. However, it still hasn't solved the problem of passing around the instance along with the function for C++ members...

    boost::bind

    What boost::bind does, is "bind" parameters into a boost::function object (how this actually happens is a bit beyond the scope of this article so I won't go into it )... like an object instance (or a pointer to one). Observe and rejoice

    class FooClass {
    public:
         void Print( int a ) {
             std::cout << "A FooClass, param = "<< a <<" this = " << this << std::endl;
         }
    };
    
    void main() {
        FooClass *myFoo = new FooClass();
        void( FooClass::* oldFunc )(int) = &FooClass::Print; //C style function pointer
        (myFoo->*oldFunc)( 5 );
    
        boost::function<void(int)> newFunc = boost::bind( &FooClass::Print, myFoo, _1 ); //boost function      
        newFunc( 5 );
    }

    What is effectively happening, is that myFoo is being "bound" into the newFunc object. Think of it as creating a private variable inside newFunc and sticking myFoo in it. When newFunc is invoked, it will use myFoo as a parameter.

    Boost is smart enough to figure out that we're passing in an instance of FooClass instead of just a number or string or whatever, and so when the function is called, it will use that instance, and will do myFoo->Print() automagically for you. This means we don't have to carry the instance round or worry about the awful syntax when we want to use it, we just bind it into the function object, and away we go.

    But but but, what's the _1?

    Aha! Remember, boost::bind is not "mysterious way to get functions to work", it's "bind a parameter into a function object." It can bind any variable of any type, so long as it matches the boost::function which will hold it. Also, it wants to bind variables for every parameter in the function. This is also perfectly valid code:

    boost::function<void()> newFunc = boost::bind( &FooClass::Print, instance, 6 );
    newFunc(); //this will call Print with 6 as the 'a' parameter.

    What the heck, I hear you say? We're calling the Print function, which takes one int as a parameter, but we're not passing any int's to it. This is just boost::bind doing it's thing. Remember, it wants to bind every parameter it can.

    ...But if you bind every parameter, you can't supply them later!

    This is what the _1 is for. It means "The first parameter, which will be supplied later". In our first example, we used _1 to indicate that we wouldn't bind the parameter to newFunc in straight away, and that we would supply it later on when we invoke the function.

    The cool thing about this is, you can mix and match them, and do all kinds of silly stuff, like this:

    int MessageBox( HWND, char*, char*, int ); 
    //Probably looks familiar to windows coders
    
    boost::function<int( int, char*, char*, HWND )> reversed_params =
          boost::bind( &MessageBox, _4, _3, _2, _1 ); 
          //use the _ markers to move the parameters around
    
    boost::function<int( char*, char* )> bind_some =
          boost::bind( &MessageBox, m_hWnd, _1, _2, 0 ); 
          // bind in some parameters, but leave others to be supplied later,
          // creating a function where the user must supply 2 params instead of 4 

    If you're a functional programming fanatic, you'll have been whinging about closures for years, and how langauges that don't have them suck. boost::bind isn't a closure, but if you know what you're doing you can get pretty close to acheiving the same functionality as one, which is pretty cool for C++

    Anyway. I've spent ages writing this so hopefully someone will read it. Good Luck!

    Sunday, September 17, 2006

    Windows live writer is catastrophically broken

    As with other microsoft editors, live writer decides that it will kindly reformat your HTML for you. By 'reformat' I mean "remove all the line breaks"

    Now, if you're just having a rant, like I am here, you really don't care. So long as the HTML isn't full of <p><p></p></p> like so many other WSIWYG editor outputs, or has (god forbid) css classes of -mso-x-y-random-otherthing everywhere, I don't mind

    Except though, when you're trying to write a post which contains code, and that code is inside a white-space:pre; element, like I am. Then, you care a lot about programs that delete all your newlines and turn your 6 line clearly-written code example into a gibbering mess.

    It turns out also, that if you tell live writer to download your existing blog posts from blogger, so you can spruce them up a bit with some WSIWYG loving, it removes all your newlines too, so if you dare republish them, they'll be a gibbering mess too.

    So, I ask you, windows live writer team. What the hell kind of good is the HTML editing mode if the app is just going to reformat and screw over any HTML you care to write? Why not just change the menu item to spawn a dialog which says "ALL YOUR HTML ARE BELONG TO US." and then quit the app? That at least would have saved me a bunch of time.

    Windows Live Writer part 2

    Well, it's actually _really_ good. Yesterday blogger was throwing error 500's like crazy so I guess that stopped it.

    Anyway, I'll be using live writer for blogging from now on I think, it sure beats publishing via the website textboxes, that's for sure.

    More stuff about C++ threading and other stuff coming soon...

    Friday, September 15, 2006

    This is a test of windows live writer beta

    Apparently it's quite good, but well... if it works, perhaps.

    Wednesday, September 13, 2006

    Your friendly QueueUserAPC

    I'm a reasonably active reader of programming.reddit.com and lately there has been a whole ton of articles about concurrent programming, so here's my 2c to throw into the ring. Basically, there seem to be 2 approaches to concurrent programming these days - the classic shared memory/locking model, which the majority of applications these days use, and the 'erlang style' nothing shared/message passing model. They both have their uses - I guess if you were using erlang you'd do that, if you were using C# you would use shared/locking. My job is to program mostly C++ apps on windows, so I'd like to share what I think is a neat trick you can use ( when programming C++ apps on windows ). This may be commonly known by everyone, but I've not seen it mentioned on reddit or anywhere else, so I'm guessing that it's not. In windows NT4 and up, native windows threads have what is called an APC Queue. APC stands for Asynchronous Procedure Call. This is used if you are using any of the Overlapped IO functions ( ReadFileEx and WriteFileEx amongst others ). Basically what it is, is a list of function pointer/ULONG_PTR pairs. So, you ask, what good might that be? Well, this APC Queue gets processed whenever the thread enters what is called an 'Alertable Wait State'. "Alertable Wait State" is a fancy name for "Has called the SleepEx function, WaitForSingleObjectEx function, or one of the many other Wait*Ex functions". So, wherever you'd do a WaitForSomething, you do a WaitForSomethingEx, and windows automagically pulls things off the APC queue and executes them. The QueueUserAPC function allows you to insert your own functions into this Queue. In a nutshell, it says "Execute this callback function in this thread" If you haven't clicked onto why that's so cool, bear with me... Observe the following program. If you're a windows C/C++ coder, it probably looks somewhat familiar #include "windows.h" #include <list> std::list<int> g_listOfInts; DWORD WINAPI ThreadProc( LPVOID param ) { g_listOfInts.push_back( 7 ); } void AddToList( int param ) { g_listOfInts.push_back( param ); } void PrintList() { std::list<int>::iterator iter; for( iter = g_listOfInts.begin(); iter != g_listOfInts.end(); ++iter ) std::cout << *iter << std::endl; } void main() { DWORD dwThreadId; HANDLE hSecondThread = CreateThread( NULL, 0, ThreadProc, 0, &dwThreadId ); AddToList( 5 ); //... do some other stuff PrintList(); } Now, this code of course has a big nasty race condition. As both threads insert into the list, they could both do it at the same time, which would either cause the list to be invalid, the program to crash, or memory corruption, or any number of other bad things. Also, the second thread could modify the list while the iterator is looping over it, which is also problematic. The classic solution is to lock the list. You could use a windows CRITICAL_SECTION, a boost::mutex::scoped_lock, or dozens of other things, which all boil down to "If any other thread wants to look at this object, it must wait for any other threads which might also be lookin at or modifying it. Also, if everything that needs to access the object must wait for everything else, we effectively serialise access to that variable down to 1 thread - if we've got 32 CPU's running 32 threads, 31 of them are going to be waiting on our lock, so we have zero performance improvement over just running a single thread. The 'non-shared' solution would be to somehow enforce that one thread "owns" the list. If any other threads want to get any data from it, they must send a message to the "owner" thread, and it must reply. This is probably trivial in something like erlang, but I don't know erlang, so I can't comment. In Windows/C++, you've got trusty old Windows Messages ( using MSG and PEEKMESSAGE, etc, like you'd have in any windows GUI app ). However, if you were to use this approach for any nontrivial program you'd end up creating hundreds of GET_X and GET_Y messages, and it would soon become unmanagable. QueueUserAPC to the rescue! Look at the next program. #include "windows.h" #include <list> std::list<int> g_listOfInts; HANDLE g_terminateSignal; DWORD WINAPI ThreadProc( LPVOID param ) { g_listOfInts.push_back( 7 ); while( WaitForSingleObjectEx( g_terminateSignal, INFINITE, TRUE ) == WAIT_IO_COMPLETION ); //apc's can execute in this loop while we wait for the quit signal. } void CALLBACK ApcAddToList( ULONG_PTR param ) { g_listOfInts.push_back( (int)param ); } void CALLBACK ApcPrintList( ULONG_PTR param ) { std::list<int>::iterator iter; for( iter = g_listOfInts.begin(); iter != g_listOfInts.end(); ++iter ) std::cout << *iter << std::endl; } void main() { g_terminateSignal = CreateEvent( NULL, TRUE, FALSE, NULL ); DWORD dwThreadId; HANDLE hSecondThread = CreateThread( NULL, 0, ThreadProc, 0, &dwThreadId ); QueueUserAPC( ApcAddToList, hSecondThread, 5 ); QueueUserAPC( ApcPrintList, hSecondThread, NULL ); //magically assume we've written some code to wait for a WM_QUIT //and set g_terminateSignal when we get it. } So, what's the difference here (apart from the code looking all werid and different)? Well, we create our second thread, which adds something to the list, like it did last time, but instead of exiting, it does a WaitEx for the terminate signal to be set. It's now in the "Alertable Wait State". Also, our main function doesn't mess with the list any more. The program is written so the second thread "owns" the list. If the main thread wants to modify it, he must make it happen in the second thread. In this example, first ApcAddToList and then ApcPrintList are "Queued" to the thread (which is in it's alertable wait state), where they are executed. Because everything involving the list only ever happens in thread 2, we no longer have our race condition. We don't have to lock at all either, instead of the threads waiting for each other so they can access the locked memory, they are free to carry on doing other stuff while thread 2 does whatever it needs to. Just like as if we'd written it in one of those fancy concurrent no-shared-state languages, but without having to rewrite your entire codebase. Cool, no? PS: If you're thinking "That's cool, but how can it be useful given that the APC callback has to be a free C-style function and only has one 32 bit parameter...", the answer will come soon. PPS: for those of you that can't wait for me to explain how it can be more useful, go and look at boost::bind.

    Friday, December 30, 2005

    New Years Eve Eve

    So, it's the 30th of december, tomorrow morning I'm going to whangamata for new years. w00t. Reading the interweb tonight, Some interesting things so I don't forget them: http://www.alfiekohn.org/articles.htm I've not heard of Alfie Kohn, but this stuff makes a lot of sense. I particularly like the 'Five reasons to stop saying 'Good Job' essay. Must read more later http://zoomin.co.nz The only map site I knew about which worked in NZ was wises nz maps. However, their interface sucks (ie: it's not google maps). Problem is now solved.

    Tuesday, December 20, 2005

    Programming Concepts

    The other day I was having a conversation with dave (my flatmate) about school and university. Anyway, he made the point that he could never just remember stuff by reading it. He said he had to understand the subject, and once he did, would then remember it naturally. I agreed with him, and so did my girlfriend. In fact, I've had the same conversation with at least half a dozen people in the past six months or so. Seeing as I'd also been reading lots of programming related stuff about which programming language is the best, the benefits of typing systems and all kinds of other geekery, this got me thinking about when I was learning to program, and how (with the benefit of hindsight) to apply that. I currently know how to program, because I understand it. But how? They say that the best way to really learn something is to teach it, so why not, I hrmmmed to myself, write a bunch of joelonsoftware/paulgraham style essays explaining how I understand programming and programming languages?

    I don't want to beat around the bush

    So, a blog. Why? Everyone else is doing it! To be honest though, I've been reading lots of joelonsoftware.com and paulgraham.com and all the 8 trillion other blogs which link off of those ones or are otherwise similar. Especially the ones on reddit, as reddit has basically become my replacement for slashdot/other tech ranting. I'll add my 5c to the mix and if nobody ever sees it well at least I can look back tenderly upon my ignorance in 10 years and cringe :-)