Showing posts with label programming. Show all posts
Showing posts with label programming. Show all posts

Saturday, July 29, 2017

JUnit Test for Singletons/Static Variables

JUnit runs all tests in one instance of JVM. Therefore, it assumes individual test cases are independent and should not interfere each other. However, this is not true for singletons or static variables. The status of singletons and static variables carry over across test cases. This causes troubles for JUnit tests.
Let me show you an example. Assume we have a MyService class: MyService class is a typical implementation of singleton. In addition, it has two methods:
init(String config)
to initialize MyService with a config parameter. This method can only call once.
doSomeService(FileWriter)
ask MyService to do something. In this example, it writes the value of someConfig via FileWriter parameter.
Now, we want to implement some JUnit test classes to verify the implementation. The first test class initializes MyService with "Hello!" and verifies whether doSomeService(...) really tries to write "Hello!". The second test class initializes MyService with "Hi!" and verifies whether doSomeService(...) really tries to write "Hi!". If you run the above test cases one-by-one in your IDE, it works fine. However, if you run both test cases in one go, one of the test cases fails:
As mentioned at the beginning, JUnit tests are run in one instance of JVM. If you run both test cases, because of singleton design pattern, the MyService instance of first test case is carried over to the second test case. When the second test case calls the init(...) method, it throws exception. This is actually one of the infamous disadvantages of singleton design pattern. However, you may not be able to re-design your system. Then, how can we fix the unit tests?

Actually, there is a workaround to force reloading of classes for each test using class loader trick. For each JUnit test class, we use a customized test runner which enforce a separated instance of class loader for each test class. Then, the classes we want to test will be reloaded for each JUnit test class Here is the implementation of the test runner (SeparateClassloaderTestRunner):
To use the SeparateClassloaderTestRunner, you need to replace the "@RunWith(MockitoJUnitRunner.class)" to "@RunWith(SeparateClassloaderTestRunner.class)" in the test classes. Then, if you run both test cases in one go, both test cases will pass.
You can get the demo source code at github - https://github.com/loklam/reloadclassdemo.

Thursday, July 27, 2017

Working Copy - a good Git Client for iPad

Recently, I want to read some source code located in GitHub. I found there is a handy tool - Working Copy. It's a Git Client for iPad. It has a nice code browser and code viewer/editor with syntax highlight. You can even edit the code and commit the changes. Now, I can read code during travelling.

Sunday, May 14, 2017

a lightweight code generator - cog


Recently, I need to write a library to read and write JSON messages for C++. The library should be able to serialize/deserialize an message object to/from a JSON message. However, the messages schema can be changed from time to time for new requirements. So, I need to think a way to adapt the schema change easily.

In Java world, there are many mature JSON libraries support this, e.g. JSON-B, Gson, .... However, C++ doesn't have reflection. We cannot have C++ version of libraries like JSON-B and Gson. For C++, the only option is code generation. There are number of ways to do code generation. After some search on the web, I found a Python based code generator -- Cog, and here is a success story about using Cog. The tool works very good and it's lightweight. With the code generation, it saved me a lot of time when there is any schema change.

Monday, January 02, 2017

jq - a powerful unix command line tool for handling JSON data

Nowadays, JSON is a popular data format, especially in web services domain. However, traditionally, Unix command line tools, e.g. sed, awk, grep, etc., are not capable to handle JSON data format. So, it is hard to write shell scripts to reading JSON data.
Fortunately, I found a handy and powerful tool - jq. In this page, I want to show you an example of using jq to extract weather forecast in JSON format to CSV. First of all, I get the data from weather forecast serveice of api.openweathermap.org and save it to "forcecast.json" file:
Here is the raw data, hard to read by human,
jq can print the JSON in pretty format, with color:
Now, I show you how to extract the dt_txt(forecast time), temp (forecast temperature) and humidity (forecast humidity) from the JSON data to CSV:
Just a  single command can covert JSON to CSV. How powerful is it! For the usage details of jq, you can refer to the manual. 

Saturday, November 26, 2016

Theory vs Practice -- Linked List and Array

When I was a student, in algorithm course about data structures, we were comparing arrays and linked lists. Arrays are faster than linked lists on random access (O(1) vs O(n)). However, listed lists are faster than arrays for random insertion/deletion operations (O(1) vs O(n)).

In reality, when we consider the L1 and L2 cache in a computer system, the above is not always true. Recently, I found a good page comparing the performance of C++ std::vector (dynamic array), std::list (linked list) and std::deque. When the size of an element is small, vectors and deques perform better than lists. It's because the spatial locality. In arrays, elements are packed next to each others in memory, but in lists, elements are located randomly and linked by pointers. Therefore, when iterating the elements in lists, it will have much higher chance of cache missed. In other words, lists can't get the benefit from cache.

Saturday, October 01, 2016

Correct Way to Convert Local Time with DST to GMT/UTC in C

The earth is round (not flat). Timezone and light saving handling is rocket science. Below is a program in C to demonstrate how to convert a string of local time in 'yyyy-mm-dd HH:SS' format to GMT/UTC time. Please note that line "27" is very critical. Without setting the tm_isdst to -1, DST won't be correctly converted.



 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
#include <stdio.h>
#include <time.h>
#include <string.h>

//input should be string of local time in 'yyyy-mm-dd HH:SS' format
int printgmttime(const char * input)
{
    struct tm localtm;
    struct tm gmttm;
    time_t timestamp;
    int rtval;

    memset(&localtm, 0, sizeof(struct tm));

    rtval = sscanf(input, "%4d-%2d-%2d %2d:%2d",
            &localtm.tm_year, &localtm.tm_mon, &localtm.tm_mday,
            &localtm.tm_hour, &localtm.tm_min);

    if (rtval < 5)
    {
        printf("unable to parse date time '%s'\n", input);
        return 1;
    }

    localtm.tm_year -= 1900; //tm_year = Year - 1900 
    localtm.tm_mon -= 1; //Month (0-11)
    localtm.tm_isdst= -1; //-1 means using system timezone info

    timestamp = timelocal(&localtm);

    gmtime_r(&timestamp, &gmttm);

    printf("GMT Date Time: %04d-%02d-%02d %02d:%02d\n",
            gmttm.tm_year+1900, gmttm.tm_mon +1, gmttm.tm_mday,
            gmttm.tm_hour, gmttm.tm_min);
    return 0;
}

int main(int argc, char** argv)
{
    if (argc <= 1)
    {
        printf("Wrong number of argument\n");
        printf("Syntax:\n");
        printf("\t%s 'yyyy-mm-dd HH:SS'\n", argv[0]);
        return 1;
    }

    return printgmttime(argv[1]);
}

Sunday, July 01, 2007

Boost C++ Libraries

Recently, I checked some job ad. on the web and found some jobs require knowledge about "Boost". I have not heard "Boost" before. After some googling, I visit the web site of Boost. Boost is really a great thing. It is a set of C++ libraries contains many useful programming constructs with good documentation, e.g. smart pointers, regular expression, object pools, state machines.... Lastly, it is open source and free.

Sunday, June 17, 2007

Dig out dead looping thread

It is not easy to write multi-threading applications. One of the common bugs is dead looping of a thread. To kill this kind of bugs, the first step is to find out which thread causes dead looping. However, an application may have dozens of threads. How to dig out the dead looping thread? My trick is to issue the ps -eLf command to list the information of threads in the whole system. The "time" column of the output of ps shows the CPU time have spent by the threads. Most likely, the dead looping thread would be the thread spending most CPU time. Then, drop down the LWP number of the thread. Next, you use gdb to attach the process and to debug the thread.

Saturday, May 12, 2007

Undefined symbols in C++

In C++ programming, we sometimes encounter "undefined symbols" problem during compilation or dlopen. The name of undefined symbols looks obfuscated. E.g.:
Unable to dlopen(test.so): test.so: Undefined symbol "_ZN6moduleD2Ev"
You may wonder why the symbol looks so ugly. Actually, this conversion of symbol is called name mangling. C++ supports polymorphism, this means functions can have same name but different types and numbers of parameters. Therefore, compiler cannot just use the function name as the symbol. Instead, both function name and parameter types should be included in symbol naming. Name mangling is the technique to encode a function name and parameter types into one symbol.

To translate the mangled symbols to more meaningful text, we can use the c++flit utility. E.g.
ahlam@oxygen:~$ c++filt _ZN6moduleD2Ev
module::~module()
Now, you know "_ZN6moduleD2Ev" is the destructor of class module.

Sunday, April 15, 2007

Optimistic memory allocation strategy in Linux

If you know C programming language, you must know what is malloc. malloc is for dynamic memory allocation. In colleges, we learned that malloc should return NULL in case of out-of-memory. However, it is not the case in Linux. By default, Linux uses optimistic memory allocation strategy. Under this strategy, Linux assumes there always exists free memory. The memory region returns by malloc is not actually allocated until the process touches the memory region. This means the memory region returns by malloc may not be available. In case of out-of-memory, the OOM Killer in Linux will pick up one or more process to kill. This sounds strange!

Reference:
man malloc
http://linux-mm.org/OOM_Killer

Saturday, March 10, 2007

Be careful with STL strings

Please refer to the following C++ code fragment:
    char charArray[4]={'a','b','c',0,};
string str1(charArray);
string str2;
str2.append(charArray, 4);
Use the following lines to print out the contents of str1 and str2:
    cout << "str1 [" << str1 << ']' << endl;
cout << "str2 [" << str2 << ']' << endl;
The output would be:
str1 [abc]
str2 [abc]
Two strings looks same. However, does str1 equal to str2?
    cout << (str1 == str2? "equal" : "not equal") << endl;
The output:
not equal
Why not!? Let's print out the size of the strings:
    cout << "str1 size = " << str1.size() << endl;
cout << "str2 size = " << str2.size() << endl;
The output:
str1 size = 3
str2 size = 4
In short, we should be careful about the == operator of strings. It does not only compare the contents of strings, but it also compares the size of strings.

Saturday, March 03, 2007

Does it really need to Lock?

In multi-threading environment, we always face the problem of race condition -- concurrent accessing sharing data among threads. To solve the problem, we can use locking primitives, e.g. mutex, to avoid concurrent accessing of sharing data. However, those locking primitives are expensive, since they involve system calls. In some cases, we can avoid using locks.

Counting
Suppose threads updating a count variable concurrently. To avoid race condition, I saw some implementation like:
mutex.lock();
count++;
mutex.unlock();
Increment an integer just takes one CPU instruction, but locking and unlocking of mutex takes hundreds or thousands of CPU instructions. To avoid race condition, we can use atomic operations provided by CPU. Referring to /usr/include/asm-i386, there is implementation of atomic add:
static __inline__ void atomic_add(int i, atomic_t *v)
{
__asm__ __volatile__(
LOCK_PREFIX "addl %1,%0"
:"=m" (v->counter)
:"ir" (i), "m" (v->counter));
}
In Apache Portable Runtime project, it provides a set of atomic operations for different platforms.

Circular Buffers
In general, without locking, a circular buffer cannot be thread-safe. However, under a restricted condition and implementation, there would be no race condition problem. In short, for a fixed size circular buffer, if there is exactly one reader and one writer, it does not need locking. It is because the reader only updates the read-pointer and the writer only updates the write-pointer. This issue have been discussed in http://ddj.com/dept/cpp/184401814.

Saturday, January 27, 2007

Stack Size of Threads

Several months ago, my boss assigned me to do performance tuning on several applications. One of them was a server application, which was for distributing data to clients. This application used a lot of threads, two threads per client -.-" It can support up to about 80 clients. If the number was over 80, the application generated a core dump. The core dump was due to out of memory. From top, I found the virtual memory size increases 4x MB for each newly connected client. It confused me. How could one client eat 4xMB?
Finally, I found the answer -- stack size of a thread. In the Linux system, the default stack size of a thread (in case of pthread library) is set to 20MB!
To change the stack size of a thread,
  1. use pthread_attr_setstacksize (&attr, stacksize) to set the attribute during thread creation. For more details, you can refer to the following link: http://www.llnl.gov/computing/tutorials/pthreads/#Stack
  2. use ulimit -s nnnn command to change the default stack size of pthread, where nnnn is the size in KBytes. (http://kbase.redhat.com/faq/FAQ_43_8710.shtm)

Monday, January 01, 2007

Bufferring Behaviour of stdout(Stardard Out)

Recently, I encounter a problem of capturing stdout of a program. The program is expected to run for a long time. When the output of the program is on the console, everything works fine. However, when the output of the program is redirected to a pipe or a file, the output is buffered for a long time. This means the output of the program is not appeared immediately. This behaviour is not desired. To illustrate the problem, I write the following simple program:
int main(int argc, char** argv)
{
while (1)
{
sleep(1);
printf("hello!\n");
}
}
The program prints "hello!" for every second. However, the output will be buffered if it is
redirected to a pipe, like:
a.out | tee tmp.log
After some investigation, I found the behaviour is documented in setbuf(3) man page:
The three types of buffering available are unbuffered, block buffered, and line buffered. When an output stream is unbuffered, information appears on the destination file or terminal as soon as written; when it is block buffered many characters are saved up and written as a block; when it is line buffered characters are saved up until a newline is output or input is read from any stream attached to a terminal device (typically stdin).... Normally all files are block buffered. When the first I/O operation occurs on a file, malloc(3) is called, and a buffer is obtained. If a stream refers to a terminal (as stdout normally does) it is line buffered. The standard error stream stderr is always unbuffered by default....

There are two ways to solve the problem. First, call fflush(stdout) after printf.
Second, call setlinebuf(stdout) at the startup of the program.

Thursday, December 21, 2006

C Macro Tricks

If you know C programming, you should know C macro as well. Macro is very powerful. However, it also makes the code difficult to read, especially macro with multi-level expansion. In complication of source code, the macro expansion is handled by "C Preprocessor". To do trouble-shooting on macro expansion, we can use the "-E" option in gcc. The "-E" option tells the compiler to stop after "C Preporcessor". With "-E" option the output file of gcc will be a text file containing the source code with all macro expanded and #include files merged.

Thursday, December 14, 2006

gdb stops at SIGPIPE

By default, gdb captures SIGPIPE of a process and pauses it. However, some program ignores SIGPIPE. So, the default behavour of gdb is not desired when debugging those program. To avoid gdb stopping in SIGPIPE, use the folloing command in gdb:
handle SIGPIPE nostop noprint pass

Friday, November 24, 2006

STL list size() method is slow

In STL list, the size() method is o(n), where n is number of elements in the list. The implementation of size() is by traversing the linked list and counting the nodes one by one. So, it sounds stupid.
I have written a simple program to test the performance of size() method in STL list and STL vector. For a STL list with 10M integers, it takes 0.17 sec. to get the size. However, for a STL vector with 10M integers, it takes 0.4 micro sec to get the size() in same machine.
There are some suggestions:
  1. Use vector instead of list.
  2. If the application need to check whether the list is empty or not, uses "list.emtpy()" instead of "list.size() != 0".
  3. use an extra counter variable to counting the size of a list.

Wednesday, December 14, 2005

CSV File Format

CSV stands for Comma Separated Values. It is a common file format for storing tabular data, e.g. spread sheet data. The file format of CSV is simple -- a text file, values are separated by comma (,) and rows are separated by newline. However, there is still some tricky in the format.

The tuck point is how to escape the comma in the values. If a value contains commas, the value should be quoted with double-quote ("). e.g.:

value one,"value two with ',' inside",value three
If a value contains a double-quote, the value should be quoted with double-quote and the double-quote in the value should be escaped with another double-quote. e.g.:
value one,"value two with '""' inside",value three
The specification of the CSV is described in RFC4180.

Even there is an RFC standard for CSV format, there are some deviation standard exists. In some applications, the comma and double-quote is escaped by a back-slash (\). Moreover, the character encoding is not stored in the file format and causes ambiguity. So, it is not trivial while using this format.

Monday, October 31, 2005

Java Generics

J2SE 5.0 have been released for a long time. However, I didn't have time to study it. Today, I spend some time to have a look of it's new features. One new feature introduced in J2SE 5.0 is generics.

In Wikipedia, the term generics is defined as:

generics is a technique that allows one value to take different datatypes (so-called polymorphism) as long as certain contracts such as subtypes and signature are kept. The programming style emphasizing use of this technique is called generic.

For example, a List object can contains different elements String, Integer, etc. In generics, we can declare a List object only contains certain types in elements. E.g, List as a List object only contains String objects.

Generics is easily mixing up with the term template. In Java, the syntax of generics is similar to Template in C++. However, generics and Template are two different concept. Template is about code generation, and generics is about type checking. You can refer to a Wikipedia's article: Comparison of generics to templates.

With generics, it can avoid type-safe problem of collection framework of Java. However, the generics in Java is very complicated. It is not elegant. After you read the tutorial of generics, you may feel very confusing, and you may find that there are many exceptional cases need to take care when using generics in Java.

Tuesday, October 25, 2005

Makefile to Build Homepage

After study GNU Make yesterday, I write a generic Makefile for building my homepage today.

The the following is the Generic Makefile:

PHP=php

all: premake all_subdirs all_curdirpages postmake

premake:
if test -x ./premake.sh; then ./premake.sh; fi

postmake:
if test -x ./postmake.sh; then ./postmake.sh; fi

include Makefile.dep

all_subdirs: $(subdirs)
for dir in $(subdirs); do (cd $$dir; make); done

all_curdirpages: $(pages)

$(filter %.html,$(pages)): %.html: %.tpl
$(PHP) $< > $@
touch timestamp

clean:
for dir in $(subdirs); do (cd $$dir; make clean); done
rm -f *.html timestamp *.gen

.PHONY: all clean all_subdirs all_curdirpages premake postmake $(subdirs) $(dep_phony)

This Makefile is put (symbolic-link) in all sub-directories of the homepage. In other words, all sub-directories use the same Makefile. So, it is called generic Makefile. If it needs to add some extra logics before or after make, we can create "premake.sh" and "postmake.sh" shell scripts respectively. The scripts will be executed accordingly. In the Makefile, the marco "pages", "subdirs" and "dep_phony" are provided by the "Makefile.dep" The customization depends on the "Makefile.dep". In each directory, it should contain a "Makefile.dep". The following is a sample "Makefile.dep":

subdirs=blog secImage

pages=index.html tba.html contact.html

index.html: index.tpl header.tpl footer.tpl rightbar.tpl \
lib/search.tpl lib/acknowledgement.tpl blog/recent/index.html

contact.html: contact.tpl header.tpl footer.tpl rightbar.tpl \
lib/search.tpl lib/acknowledgement.tpl

tba.html: tba.tpl header.tpl footer.tpl

The "subdirs" macro defines the the sub-directories requiring to process, and the "pages" macro defines the HTML pages requiring to generated. The "Makefile.dep" file also defines the dependency of HTML files to the template files (*.tpl).

To build the homepage, I only need to issue the "make" command in the root directory of the homepage source.