Tuesday, July 10, 2007

The real jungle

Just back from teaching Perl in Singapore. The real highlight for me, apart from teaching Perl of course, was the view of the Malaysian jungle on the flight in. The area looks tiny on the map, but it is vast. An amazing place.

I found Singapore to be hot and steamy, and I am talking about the climate. The place has still not shaken off its colonial past. Its own character does come through sometimes, and more power to it. The people seem to be friendly and open without the hang-ups of many large cities. Of course I was only there for a week and I only saw a small part of it.

As usual I looked at scripts that delegates brought in. I have also had a couple of scripts from other people recently. I see very common bad practices regularly:

No use warnings
No use strict
All production scripts should have these set.

Calling subroutines using the & prefix. Why do people do this? Because that is the way we used to do it in perl 4, and people do not update their skills. The & prefix ignores prototype checking, and passes the current value of @_ into the call if no other argument is specified. Don't do it!

Calling external programs instead of Perl built-ins. I have seen `date` used instead of localtime(), `cd $dir`, `pwd`, `mkdir`, `rm`, `egrep`, and so on. This is not just lazy, it is grossly inefficient. It is also a security risk, by the way. I published a list of UNIX commands and perl equivalents, I will try and update it soon.

Monday, June 25, 2007

The threading jungle

Multi-threading has been around a long time. I first heard the term in the mid-1980s when discussing how ICL were going to host UNIX on VME. The first real implementation I saw was on Apollo NCS, a system that had many innovative features. ICL, VME, and Apollo have all gone, victims of acquisitions. Threads remain.

Multi-threading hit the mainstream with Windows NT, the system was built on it. The consensus around us UNIX hackers was that Microsoft had to encourage threads because their processes were so inefficient - no fork/exec you see. Probably utter twaddle; you know what UNIX hackers are.

Threads brought synchronisation problems. Of course they did. System programmers were no strangers to this type of issue, just the context of threads was new(ish). Way back when, on ICL VME, all we had were "test and set" and "decrement and set" instructions that were atomic. When I were a lad we had to build our own primitives, there was a particularly handy instruction that let you not only raise an interrupt (event/signal, whatever) but allowed you to send a two-word data block with it. Kids of today … don't know they're born… forty miles 't pit… etc.

I got to write real synchronisation primitives when I did the VME port of the MIMER RDBMS in the early '80s. There is nothing like having to write the damn things to understand how they work. I did not design them, I hastily add - that was done by some egg-head at Uppsala University. The problem then, as now, is not in using the APIs, it is in understanding how they interact with your application and (more important) how the application interacts with itself.

First rule when writing co-operating threads: things only go wrong at the worse possible moment. And that is usually after the thing has gone into production.

The Windows APIs are fairly easy to use, certainly compared to knitting your own. The main wrinkles in using them are a few peculiarities with the C/C++ runtime-library. I find it amazing that so many system engineers insist in using CreateThread instead of _beginthreadex. Back in Visual Studio 2.0 this was understandable, even I did it. The documentation was, um, "challenged". Now, with the MSDN, there is no excuse, the problems are well documented. Still people insist on using the wrong API. My theory, for what it is worth, is that a call to _beginthreadex, with its associated casts, looks messy on the screen, surrounded as it is by unrivalled elegance (sic). CreateThread looks good, _beginthreadex is ugly. For wotsit's sake! A thing of beauty with a memory leak is bad code, and I don't care what it looks like. Humph.

Still, you thought that was ugly? Try finding the thread-safe RTL functions in pthreads. Threads had to be retro-fitted to UNIX. Many fought against it, even the hero Torvold in the clone wars, but resistance was futile. Microsoft made life easy for us developers by having a whole special runtime-library just for multi-threading, ungrateful pups that we are. The pthreads implementations on UNIX have no such luxury. Instead there is a whole mess of different functions that are re-entrant with the _t suffix. Bye-bye portability. Finding which functions these are is a bit hit and miss, sometimes they are not in the man pages, and the standard only lists minimum requirements. Try teaching Condition Variables to a class already spaced out by mutxes. The word "predicate" makes their eyes glaze over. The fixed grin is the give-away: "Gone, solid gone", as Baloo would say. I'm sure Kipling was a coder, who else would write a whole poem about a conditional statement?


Wednesday, May 30, 2007

What I did on my holiday

1. Read "Perl Hacks" again, and caught a bunch of stuff I had missed first time around. Like, wantarray returns undef in void context.

2. Read "Parallel Programming in C with MPI and OpenMP" by Michael J. Quinn
Very American (the first computer was ENIAC - really?) No mention of the ICL DAP (Distributed Array Processor) - one of the first commercially available parallel processors.
Nice to see terms like "Functional Decomposition" from the 1970s recycled to be used in a parallel computing methodology.
It really bugs me when authors come up with "Creating Arrays at Run Time" and give code like this:
int *A;
A = (int *) malloc (n * sizeof(int));
When there is a perfectly good feature in C99 to do it:
int A[n];
That's it! This is a C99 Variable Length Array, one of the more obvious features of the changes from C89. The book was published in 2004, so why no mention?
The author uses for loops to access arrays quite often, but there is no mention of optimisers, and the effect of optimisations on parallel processing. Ho hum.

Thursday, April 05, 2007

Where have you been?

Guilty. I haven't blogged for a long time.

What's happened?
Well for starters the company I work for got taken over, and that caused a lot of dust - mostly settled now.

I have been getting into the new Korn shell 93, and I hope to post something on that soon.

Just today I got my "Happy Monkday" from perlmonks - I have been a monk for a year. Good fun!

Oh where is Perl 5.10? Soon? Please!

Thursday, August 31, 2006

Notifications

This is a reminder post for (selfishly) me! Follow along if you like.
(CPAN) Win32::ChangeNotify - does it support (Win32 API) ReadDirectoryChangesW and overlapped IO? If not, how does it avoid missing changes?

(CPAN)Linux::Inotify (and Inotify2) support the Linux kernel API inotify which (perhaps) replaces dnotify. Supported from kernel 2.6.13.
Neither interface is on Fedora core 4, which is 2.6.11.

Further digging required.

Monday, July 24, 2006

Fun with Perl module loading

Q. So exactly when does a BEGIN block get executed?

A. As soon as it is encountered!

For example, take the following code:

use AModule;
use BModule;

BEGIN {
print __PACKAGE__." BEGIN block\n";
}

use AnOther;

$syntax_error = 42

print "Starting main program\n";
print "Ending main program\n";

The order of execution is:
   AModule BEGIN block
AModule main body

BModule BEGIN block
BModule main body

main BEGIN block

AnOther BEGIN block
AnOther main body

syntax error at main.pl line 18, near "print"
Execution of main.pl aborted due to compilation errors.

So, the main program's BEGIN block is not necessarily the last one executed. Granted, normally it is, since we usually use all our modules before the BEGIN block in main, but we don't have to.
Note also that the blocks are executed even though we have a syntax error, and the program fails to compile.


Q. What use is that?

A. We can alter the way subsequent modules are loaded

Usually that will be by altering @INC:

BEGIN {
if ( defined $ENV{TESTLIB} ) {
unshift @INC, $ENV{TESTLIB}
}
}


Now all subsequent module loads will search the directory in the environment variable first.

Q. @INC is just a list of directories, Right?

A. Wrong! It can also contain code references.

References to subroutines found by the module loader will be executed as they are found:


BEGIN {
print __PACKAGE__." BEGIN block\n";
my $code = sub { print "\@INC code\n" };
unshift @INC,$code;
}
use AnOther;

Gives:

main BEGIN block
@INC code
AnOther BEGIN block

Q. Is that it?

A. Of course not.

The subroutine is passed two arguments, the code reference itself and the name of the module it is trying to load. This enables us to track every module loaded. So:

BEGIN {
print __PACKAGE__." BEGIN block\n";
my $code = sub {
my (undef, $loading) = @_;
my ($package, $filename, $line) = caller;

print "Loading $loading from $package\n"
};
unshift @INC,$code;
}

use AModule;
use BModule;
use AnOther;

Gives:

main BEGIN block
Loading AModule.pm from main
AModule BEGIN block
AModule main body

Loading BModule.pm from main
BModule BEGIN block
BModule main body

Loading AnOther.pm from main
AnOther BEGIN block<>
Of course, further diagnostics could be added, like filename, line, and date/time. I'll leave that to you.

Q. So, @INC is more powerful than its sister %INC, which is just used for diagnostics.

A. Err, no. %INC is used by perl to see if a module is already loaded.

Q. And how is that useful?

A. We can force a module reload by removing its entry from %INC. Consider:
In one part of our code we want to force a different version of a module to be loaded (don't ask why).
  use strict;
use warnings; # DON'T use -w

use A;
use MyB;
use C;

A::mysub(); # Original modules used

delete $INC{'A.pm'}; # Force perl to reload

unshift @INC,'mydir'; # Change @INC
{
no warnings 'redefine'; # No 'redefined' warnings
require 'A.pm'; # Reload module
}

A::mysub(); # Module in 'mydir' used

Work it out yourself!

Thursday, May 04, 2006

Splitting whitespace

The documentation for Perl is quite good usually, but when it comes to perldoc –f split there is so much magic involved that normal English begins to break down. Take the default arguments, for example.

"If EXPR is omitted, splits the $_ string. If PATTERN is also omitted, splits on whitespace (after skipping any leading whitespace)."

Then, later in the documentation:

"As a special case, specifying a PATTERN of space (' ') will split on white space just as "split" with no arguments does. … A "split" with no arguments really does a "split(' ', $_)" internally."

So, how many whitespace characters is that? A single space as a delimiter implies a single space is used for the split, but it actually does 'one or more whitespace':
   $_ = '    This    is    some    text';
@a = split;
$" = '|';
print "@a\n";
Produces
   This|is|some|text
So leading whitespace is ignored, and one or more whitespace is used as a delimiter. There is a (documented) subtle difference with \s+:
   $_ = '    This    is    some    text';
@a = split /\s+/;
$" = '|';
print "@a\n";
Produces:
   |This|is|some|text
Notice that the first element of the resulting list is empty, which was not previously the case.

A few questions arise from this. First, what is this ' ' syntax all about? Don't we need a regular expression match?
   $_ = 'xxxThisxxisxxsomexxtext';
@a = split 'x';
$" = '|';
print "@a\n";
Produces:
    |||This||is||some||te|t
So it does work, except not exactly the same as a single space, it does not match 'one or more' (x+), so the space is magic. To be fair the documentation does say that ' ' is a special case. But the documentation does not show the syntax of a string literal, it specifically shows that an RE delimited with / / is required. Single quotes works with regular expressions, and with multiple characters (without a leading 'm'). But double quotes or other characters do not work unless preceded with 'm'.

Second question. What does whitespace mean? Is ' ' the same as \s in this case? Normally, of course, it is not, but in this case it is! ' ' is very special.

Thursday, April 27, 2006

Inside-out accessor methods - version 2

This solution uses a hash of hash-references, which means we no longer need the nasty eval. With thanks to the perl monks.

my (%speed, %reg, %owner, %mileage);
my %hashrefs = ( speed => \%speed,
reg => \%reg,
owner => \%owner,
mileage => \%mileage);

sub set {
my ($self, $attr, $value) = @_;
my $key = refaddr $self;

if ( !exists $hashrefs{$attr} ) {
carp "Invalid attribute name $attr";
}
else {
$hashrefs{$attr}{$key} = $value;
}
}

Wednesday, April 26, 2006

UNIX commands to Perl

This is a rough conversion table from UNIX commands to Perl.
Please note that there are not always direct single equivalents.

UNIX Perl Origin
. do built-in
awk perl ;-) (often 'split') built-in
basename File::Basename::basename Base module
cat while(<>){print} built-in
ExtUtils::Command::cat Base module
cd chdir built-in
chmod chmod built-in
chown chown built-in
cp File::Copy Base module
ExtUtils::Command::cp Base module
date localtime built-in
POSIX::strftime Base module
declare see typedef
df Filesys::Df CPAN
diff File::Compare Base module
dirname File::Basename::dirname Base modules
echo print built-in
egrep while(<>){print if /re/} built-in
eval eval built-in
exec exec built-in
pipe (co-processes) built-in
export Assign to %ENV Hash variable
Env::C CPAN
find File::Find Base module
ftp Net::Ftp Base module
function sub built-in
grep see egrep
integer int built-in
kill kill built-in
ln -s link built-in
ls glob built-in
opendir/readdir/closedir built-in
stat/lstat built-in
mkdir mkdir built-in
mkpath ExtUtils::Command::mkpath Base module
mv rename built-in
ExtUtils::Command::mv Base module
od ord built-in
printf built-in
print print built-in
printf printf built-in
rand rand built-in
rm unlink built-in
ExtUtils::Command::rm Base module
rm –f ExtUtils::Command::rm_rf Base module
sed s/// (usually) built-in
sleep sleep built-in
alarm built-in
sort sort built-in
source do built-in
times times built-in
touch open()/close() built-in
ExtUtils::Command::touch Base module
trap %SIG Hash
sigtrap pragma
typeset my built-in
typeset –I int built-in
typeset –l lc built-in
typeset –u uc built-in
typeset -Z sprintf built-in

Tuesday, April 11, 2006

Sad....

Walking across London I spotted a poster which screamed:
100 Gig tickets to be won!
"Must be a big stadium", I thought....

Accessor methods - update

I have a much better solution, with the help of the perlmonks. I'll post it when I get chance

Thursday, March 23, 2006

Accessor methods for inside-out objects

I have just been looking at Class::Accessor, recommended by Simon Cozens at http://www.perl.com/pub/a/2006/01/26/more_advanced_perl.html. It automagically generates accessor methods, but appears to rely on a %fields type approach, and will not work on inside-out objects.
I got to thinking. Actually generalised accessors for inside-out objects are not that difficult, not need some ev[ai]l doings.



my %speed;
my %reg;
my %owner;
my %mileage;

sub set {
my ($self, $attr, $value) = @_;
my $key = refaddr $self;

my $hashref;
eval "\$hashref = \\\%$attr";

if ( !defined $hashref ) {
carp 'Invalid attribute name';
return;
}

$hashref->{$key} = $value;
}

sub get {
my ($self, $attr, $value) = @_;
my $key = refaddr $self;

my $hashref;
eval "\$hashref = \\\%$attr";

if ( !defined $hashref ) {
carp 'Invalid attribute name';
return;
}

return $hashref->{$key};
}


BTW: I also found how to do formatting: <pre>...</pre>

Wednesday, March 08, 2006

More of the same..

I havn't posted for a while. It's not that nothing has happened, its just that too much has! I'm afraid that the blog suffers when I'm busy. I have been madly updating our 'Advanced Perl with CGI and Web Applications' course, and I am quite pleased with the result. As always I reckon I learnt at least as much as the delegates will (although maybe on more obscure subjects).
For example, did you know you can use __DIE__ and __WARN__ signal handlers to trap Carp::croak and Carp::carp calls?

I returned home from teaching in the USA last Saturday to my latest multiple order from Amazon:
Object Oriented Perl by Damian Conway. A little out of date (2000) but good nontheless, written in typical Damian style. Hacking the Advanced Perl course (two chapters on OO) and reading this book has totally changed my views on Perl as an OO language. The book does not contain anything on inside-out objects, but you can see that Damian was almost there. We now have a chapter on this simple but effective means of encapsulation in the Advanced course.

Number two book is the "Pickaxe book", Programming Ruby. About time I read it, more on that later (hopefully).

Number three is "PHP and MySQL Web Development". A huge tome which (unlike OOP and Ruby) does not look like 'fun'.
Last but not least another O'Reilly pocket book, this time the Linux one. All these pocket books, I need huge pockets. Not to mention another bookshelf....

Monday, January 16, 2006

What a refreshing change

I'm on the West coast of Ireland, Galway, teaching. I don't think anyone will be offended if I say that Galway is a little remote.
The air and the Guiness are all refreshing, but what stopped me in my tracks is the co-operation between businesses. One major international company booked me to give a course, but then said to others in Galway, "Hey, we have a technical course running this week, anyone else interested?" , and so along comes a few guys from other companies that could not justify the course on their own. Apparently this happens all the time.
These companies are major internationals acting like sensibile human beings. Now why can't that happen everywhere?

Wednesday, January 04, 2006

IFS and read

I find myself continually underestimating the Korn shell. I always thought that Bash was neat because we could read into an array with read –a, but of course we can do the same in ksh with read –A. The effect is to split up the input into fields around $IFS. But did you know we can set IFS just for a read statement? Look carefully:

IFS=','
while IFS=':' read -A line
do
end=$(( ${#line[*]} - 1 ))
if [[ ${line[$end]} != /sbin/nologin ]]
then
echo "${line[*]}"
fi
done < /etc/passwd

(display all the lines in /etc/passwd whose last field is not /sbin/nologin, changing the field delimiter from colon to comma).

The first IFS=',' sets IFS for the echo expansion and is not overridden by the IFS on the read line. That IFS only applies for the read, nothing else.

By the way, did you know that the expansion (including $*) uses the first character from IFS, so order matters? So IFS=',.:' gives different results to IFS=':,.'.

All this applies to the standard POSIX shell, as well as Korn shell and Bash, and I was alerted to these features by Classic Shell Scripting (see a previous post).

Finally, a significant bug in pdksh has been fixed in the version with Fedora core 4 (1993-12-28 q). In early versions a pipe generated a child process on each side, so piping into a while loop was not useful, since the while loop ran in a different process to the rest of the script, and so variables could not be saved. Now it all works, consider:

#!/bin/ksh

ps -ef | while read uid pid ppid c stime tty time cmd
do
if [[ $cmd == *xinetd* ]]
then
break
fi
done

echo "Pid of xinetd is: $pid"

Unfortunately this will still not work in Bash, for the same reason as it did not work in the old pdksh. In Bash we can use process substitution instead, and while we are about it we may as well use an ERE (because we can):

while read uid pid ppid c stime tty time cmd
do
if [[ $cmd =~ '^xinetd +' ]]
then
break
fi
done < <(ps -ef)

echo "Pid of xinetd is: $pid"

The man pages for the new pdksh say process substitution works in Korn shell as well, and I am told this works on Solaris, but it does not appear to work on Linux.

Friday, December 30, 2005

Toys for boys

The weather has been kind over Christmas, particularly clear skys for my *new* Meade ETX 105 - that's a telescope by the way. Not exactly the largest - only 4" and yes, size does matter, but I couldn't really justify a larger one given that it usually rains on the rare occasions when I am actually at home.
Good views of Mars in November and early December, now good views of Saturn, and various nebula like M42 and M37. OK, so what's that got to do with programming and stuff?
Well, most modern telescopes have an electronic control called a "goto" system - you get the idea. This is normally controlled by a hand-held device, but can be controlled by a PC using an Open Source object model named ASCOM - Astronomy Common Object Module.
http://ascom-standards.org/faq.html
Currently it appears to be written using COM, which of course is proprietry and realistically only runs on one of the many Microsoft operating systems. OK, you can get it to run on others, but would you want to, really?
This is one I will be looking at further.

Wednesday, December 21, 2005

Switching off terminal echo in Perl

A delegate recently asked me how to switch echo off for password prompting. I couldn't remember, I thought it was Term::ReadLine. Actually it is Term::ReadKey - as expected the Perl Cookbook reminded me. Term::ReadKey is not a base module, but is present on ActiveState and Fedora installations, and easily installed from CPAN. Here is a simple script that works on both Linux and Windows (despite what the documentation says) :

use Term::ReadKey;

print 'Password: ';
ReadMode 'noecho';
my $password = ReadLine; # Corrected
ReadMode 'normal';
chomp $password;

print "Password was: $password\n";

Another delegate said something like "I guess another module will replace characters typed with asterisks". Well, that is not so simple. After some experimentation I came up with the following script, which has to use raw terminal input and take iinto account differences between Windows and Unix/Linux:

use Term::ReadKey;
use strict;

local $| = 1;
print 'Password: ';

my $password = '';
my $retn = $^O eq 'MSWin32'?"\r":"\n";

ReadMode 'raw';

while (1)
{
my $key;
# Loop waiting for a key to be pressed
while (!defined $key)
{
$key = ReadKey -1;
}

last if $key eq $retn;
$password .= $key;

print '*';
}

ReadMode 'normal';
print "\n";

print "Password was: $password\n";

Check the Term::ReadKey documentation for the arguments. Note that I have to wait for a key to be pressed this time, and write an unbuffered '*' (local $| = 1 makes STDOUT unbuffered).

Wednesday, November 16, 2005

Perl subroutine parameter passing - again


One of the reasons Damian Conway does not like prototypes is that (his words) they do not give the results expected. Well, what results do you expect when you don't use them?

It is sometimes said that parameters are passed into Perl subroutines by copy, and some say they are passed by reference. In fact they are passed by magic, and as we all know, any technology sufficiently advanced is indistinguishable from Perl.

What do you expect to be printed?

sub mysub
{
@_ = qw(one two three)
}

@array = qw(The quick brown fox);
mysub (@array);
print "@array\n";

We get 'The quick brown fox '. The array is unchanged, so you might think the array is passed by value. The call:

mysub (qw(The quick brown fox));

also works fine, so no problem. Let's change the subroutine to use a slice instead:

sub mysub
{
@_[0,1,2] = qw(one two three)
}

Now we get 'one two three fox', the array is changed! Changing the elements changes the caller's array, changing the whole array does nothing! Now the line:

mysub (qw(The quick brown fox));

fails with "Modification of a read-only value attempted…"

This is a run-time error, whereas if you use a prototype:

sub mysub (\@)

we get a compile-time error:

"… must be an array (not list)…"

If I preferred run-time over compile-time errors I would be using the Korn shell.

Once the prototype has been specified it does not matter if we alter the caller's array in total or use a slice – the caller's code is always the same.

OK, I admit that altering @_ direct is damn awful code, and it breaks another of Damian's rules. He's not all bad ;-)

I don't believe that an inconsistency in one part of Perl (prototypes) is a justification to switch to an inconsistency in another (pass by magic). Just admit the inconsistencies and celebrate them – or move to Perl 6.






Thursday, October 20, 2005

Perl 6 at EuroOSCON

This is a brief summary of my notes from the Perl 6 and pugs sessions. I have only included the stuff that was new compared to the Perl 6 Appendix in the 'Perl Programming' course. I might have repeated myself.


use strict and use warnings are on by default;

Quotes are only required around a hash key when {} are used.

$hash{'key'} becomes %hash

Barewords are not allowed. Ever. Even for file handles.

fail{…} warn{…} die{…} throw a 'not yet implemented' type warning – the Perl 6 developers must be using that a lot ;-)

Many improvements to interpolation:

"{ executable code like a do block} "

Subroutine calls are allowed: "&mysub(args)"

Interpolate an array: "@array[]"

Interpolate a hash: "%hash{}"

sprintf is probably obsolete, printf definitely is:

print $var.as('%6.2f');

print $hash.as("%-20s: %2.6f", "\n");

Control can be made over exactly what is interpolated:

qq:c(0)/ / # don't interpolate

qq:s(1)/ / # only interpolate scalars

qq:s(1):a(1)/ / # interpolate scalars and arrays

hummmmmmm

Here document syntax is changed:

my $var = q:to /END/ # s(1):a(1):h(1) can be added

…

END

Not sure where the semi-colon goes

Ranges have a new feature:

x..^y means from x up to y-1

x^..^y means from x+1 to y-1

To open a file with an automatic chomp::

my $fh = open $filename

This can be overridden.

while (<$fh>) {…}

becomes:

while = $fh {…}

while (<>){…}

becomes

while = $ARGS{…} or while =<> {…}

=<> is known as the 'fish' operator

There is now a 'prompt' verb

Most built-ins that operator on scalars and arrays are now methods:

$var.substr(…) # returns the substring

$var .= $var.substr(…) # changes the substring

No need for Data::Dumper, instead:

$var.perl()

And there is much more, Damian could not complete his talk because of time, and it is not downloadable. I'll keep you posted as I play with pugs.

Wednesday, October 19, 2005

EuroOSCON - whatever. Is Perl 6 too complex?

I asked the main man, yer actual Larry Wall, this question. Poor guy - he is a very nice chap - thought he was just signing a book for me. There was no one else around, so I consulted the oracle (I won't use an uppercase O, you'll get the wrong idea).

Q. Will Perl 6 still be suitable for guys who just want a better language that ksh or .bat files?
A. The easy things will still be easy

Q. Will Perl 6 still be suitable for sys.admin's?
A. Hey! I'm still a sys.admin.
Me: No you are not, you are a language demi-god.
A. Shrug. Sys.Admins will still be able to use it.
Q. Are you saying that because that is what I want to hear?
A. Sys.Admins will still be able to use it.

Thank-you very much.
We also discussed training strategies for Perl 6, and even for Perl 5, and I wasn't that hard on him. Larry Wall is a very nice chap (oh, I said that).